08/25/2026 | Press release | Distributed by Public on 08/25/2026 10:36
The Aiera Research Score ranks the response quality of 32 language models on investment research queries. But fold in what each model costs (price, speed, and average tokens per task) and the model that scored #1 drops down to 29th.
Our Research Score measures one thing: how well does a model produce institutional-grade analysis, across Facts, Grounding, and Depth. At the moment, Opus 5 tops the leaderboard with a score of 87.
That score represents the model's capability. But we also need to consider how to rank a model's value.
Three numbers no capability board can show decide value:
We combined all of these factors into a single Value Score: research quality, cost per task, and latency per task. Here's what happens to the board:
Performance is the Aiera Research Score (mean of Facts, Grounding, Depth). We measured each model's average output tokens per task across all tasks in the benchmark, then turned price and speed into what you actually experience:
cost / task = tokens × output price · latency / task = tokens ÷ speed
Value = 0.5·performance + 0.3·cost + 0.2·latency (each normalized 0-100)
Performance carries the most weight on purpose; a cheap, fast model that can't do the work is worthless. Price and speed then break the ties a capability board ignores.
While Opus 5 is the best financial researcher model; it is the worst value, with the average answer costing 15¢ (up to 30x vs. other models) and taking nearly two minutes (3x the average of other models) to complete.
Best overall value: MiniMax M3
81.5 research (top-10 quality) at 0.45¢ and ~49s per task. Near-frontier work at open-weight economics.
Frontier-grade, no tax: GPT-5.6 Terra or DeepSeek V4 Pro
80+ research at a fraction of the flagship cost and half the latency; the sweet spot when quality can't slip.
Fastest capable: Nemotron 3 Ultra or Gemini 3.5 Flash
70/68 research with answers in 10-13 seconds. When a person is waiting, speed is the feature.
Value needs a quality floor. Rank purely on cost and the cheapest, fastest models like Grok 4.3, Llama 4 Scout, and gemma would top the chart. They don't, because they score under 30 on research: dividing quality you don't have by a price you barely pay is a useless recommendation. That's why performance carries half the weight, and why those models sit low despite near-zero cost.
Opus 5 really is the best on research quality. If your task is rare, high-stakes, and latency doesn't matter (a one-off memo, not a pipeline) pay for the ceiling.
The value board isn't an argument against the best model. It's a map of what the best model costs, and of how little you give up to run something cheaper and / or faster at scale. There's no single best model, only a best model for a budget, a latency target, and a quality floor.