Point11

Analyst Bench

Which model makes the best analyst?

Adam Fish

Adam Fish Co-Founder & CEO

Sep 26, 2026

Point11 helps brands win their market by understanding their customers, and the first step is understanding the brand itself. In this post, we share what we found when we gave a dozen models the work of a senior analyst: profiling 25 companies, from their size and scope to their markets and competition. We scored each model on quality, speed and cost.

The Pareto frontier

The dotted line is the Pareto frontier: the models you can't beat on intelligence and cost at the same time. Claude Opus 5.5 scored highest, but Point11 Alpha landed within the margin of error for less than half the price. That's because Alpha doesn't rely on one model; it hands each step of the work to whichever model does it best. Past Alpha, the line flattens out, so spending more buys almost nothing.

Intelligence vs. cost

Most attractive quadrantPareto line

7 arms sit on the Pareto line (Muse Glimmer, Gemini 3.5 Flash-Lite, GPT-6 Luna, Gemini 3.8 Flash, Muse Spark 1.3, Point11 Alpha, Claude Opus 5.5). Point11 Alpha is on it.

Model intelligence

Scores varied a lot. The weakest models made about five times as many factual errors as the strongest, things like a wrong revenue line or a missed market, and Point11 Alpha made the fewest of all. Each score combines two things: quality, graded blind by the top model from Anthropic, OpenAI, Google and xAI, and accuracy against public companies' annual reports.

Intelligence

Score out of 100, with 95% confidence intervals

  1. Claude Opus 5.5: 75.8
  2. Point11 Alpha: 75.1
  3. Claude Fable 5.1: 70.4
  4. GPT-6 Astra: 70.1
  5. Muse Spark 1.3: 66.3
  6. Grok 4.7: 64.8
  7. Grok 4.6: 62.9
  8. Gemini 3.8 Flash: 62.4
  9. Claude Sonnet 5: 58.4
  10. GPT-6 Luna: 57.6
  11. Gemini 3.5 Flash-Lite: 49.0
  12. Claude Haiku 4.5: 43.2
  13. Muse Glimmer: 41.8

Ranked by Index. Point11 Alpha scores 75.1 (95% CI 72.7–77.5); the best single model, Claude Opus 5.5, scores 75.8. Whiskers are 95% intervals.

Workflow speed

The strongest single models were also among the slowest. Point11 Alpha finished well ahead of them, and the only models faster still were lighter ones that scored noticeably lower. We measured speed end to end, including Point11's own processing.

Speed

Minutes per Profile, start to finish

  1. Gemini 3.5 Flash-Lite: 12
  2. Gemini 3.8 Flash: 15
  3. Muse Glimmer: 16
  4. Claude Sonnet 5: 20
  5. GPT-6 Luna: 21
  6. Claude Haiku 4.5: 22
  7. Point11 Alpha: 23
  8. Muse Spark 1.3: 24
  9. Claude Opus 5.5: 38
  10. Grok 4.6: 40
  11. Grok 4.7: 44
  12. Claude Fable 5.1: 52
  13. GPT-6 Astra: 57

Fastest first. Point11 Alpha finishes a Profile in 23 min; the fastest single model, Gemini 3.5 Flash-Lite, takes 12 min. Slowest: GPT-6 Astra. Whiskers run to the 90th percentile.

Workflow cost

Cost varied even more than quality: the most expensive model cost about 36 times as much per profile as the cheapest. Spending more only bought more intelligence up to a point, the same flattening you see on the frontier. These are the actual costs of producing each profile, not list prices, split across the pipeline's five stages.

Cost per task

USD per Profile, stacked by pipeline stage

  1. Muse Glimmer: $0.27
  2. Gemini 3.5 Flash-Lite: $0.40
  3. Claude Haiku 4.5: $0.55
  4. GPT-6 Luna: $0.66
  5. Gemini 3.8 Flash: $1.05
  6. Muse Spark 1.3: $1.60
  7. Claude Sonnet 5: $2.40
  8. Point11 Alpha: $3.40
  9. Grok 4.6: $3.90
  10. GPT-6 Astra: $5.10
  11. Claude Opus 5.5: $7.40
  12. Grok 4.7: $7.90
  13. Claude Fable 5.1: $9.60

Cheapest per Profile first. Muse Glimmer costs $0.27; Point11 Alpha costs $3.40.

Methodology

Every model ran Point11's pipeline unchanged, with the same prompts and its lab's default settings. Point11 Alpha ran that same pipeline, with each stage assigned to the model best suited to it, and faced the same judges and answer keys. The judges never knew which model wrote which profile, and no single lab's model graded alone. Keep in mind that these results measure how models do on this workflow, not how capable they are in general. We'll rerun the benchmark as new models come out.

If you have questions about our methodology, or you'd like us to benchmark your own enterprise workflows, reach out to our sales team.