Aiera Inc.

08/25/2026 | Press release | Distributed by Public on 08/25/2026 10:36

Triangulating Value with AI in Research

The Aiera Research Score ranks the response quality of 32 language models on investment research queries. But fold in what each model costs (price, speed, and average tokens per task) and the model that scored #1 drops down to 29th.

Our Research Score measures one thing: how well does a model produce institutional-grade analysis, across Facts, Grounding, and Depth. At the moment, Opus 5 tops the leaderboard with a score of 87.

That score represents the model's capability. But we also need to consider how to rank a model's value.

Three numbers no capability board can show decide value:

  1. Price per token;
  2. Speed in tokens per second, and:
  3. Quantity of tokens per task (the one everybody forgets).

We combined all of these factors into a single Value Score: research quality, cost per task, and latency per task. Here's what happens to the board:

How the Value Score works

Performance is the Aiera Research Score (mean of Facts, Grounding, Depth). We measured each model's average output tokens per task across all tasks in the benchmark, then turned price and speed into what you actually experience:

cost / task = tokens × output price · latency / task = tokens ÷ speed

Value = 0.5·performance + 0.3·cost + 0.2·latency (each normalized 0-100)

Performance carries the most weight on purpose; a cheap, fast model that can't do the work is worthless. Price and speed then break the ties a capability board ignores.

While Opus 5 is the best financial researcher model; it is the worst value, with the average answer costing 15¢ (up to 30x vs. other models) and taking nearly two minutes (3x the average of other models) to complete.

What the Token Count Reveals

  1. Verbosity tax
    The most capable models are also the most verbose. Opus 5 and Sonnet 4.6 write ~5,900 tokens per answer - more than double the non-Anthropic model . At premium output prices, that length compounds twice: it's what you're billed for and what you wait through. Opus 5's brilliance arrives at ~15¢ and 115 seconds per task. A per-token price tag hides this entirely; a per-task view makes it the headline.
  2. Terse-and-capable wins on value
    The value board is led by models that are both strong in research and economical with words. MiniMax M3 (81.5 research) and DeepSeek V4 Pro (80.1) deliver near-frontier quality research output for well under a cent per task in less than half the time; Opus 5 costs roughly 30-50× more per answer for a few points more quality. Even inside OpenAI's own lineup the split is stark: the efficient GPT-5.6 Luna lands #2 on value, while the flagship GPT-5.6 Sol (nearly tied on research) sits at #22, sunk by a wordy 4,100-token task at premium frontier prices.
  3. Latency is the lack of brevity × speed
    Time-to-answer is where verbosity and throughput collide. Nemotron 3 Ultra, Gemini 3.5 Flash, and Gemini 3.1 Pro return answers in under 15 seconds because they're both fast and concise. The Kimi models are strong and cheap but take over a minute; not from weakness, but because they're slower and wordier.

The picks

Best overall value: MiniMax M3
81.5 research (top-10 quality) at 0.45¢ and ~49s per task. Near-frontier work at open-weight economics.

Frontier-grade, no tax: GPT-5.6 Terra or DeepSeek V4 Pro
80+ research at a fraction of the flagship cost and half the latency; the sweet spot when quality can't slip.

Fastest capable: Nemotron 3 Ultra or Gemini 3.5 Flash
70/68 research with answers in 10-13 seconds. When a person is waiting, speed is the feature.

The Net Net

Value needs a quality floor. Rank purely on cost and the cheapest, fastest models like Grok 4.3, Llama 4 Scout, and gemma would top the chart. They don't, because they score under 30 on research: dividing quality you don't have by a price you barely pay is a useless recommendation. That's why performance carries half the weight, and why those models sit low despite near-zero cost.

Opus 5 really is the best on research quality. If your task is rare, high-stakes, and latency doesn't matter (a one-off memo, not a pipeline) pay for the ceiling.

The value board isn't an argument against the best model. It's a map of what the best model costs, and of how little you give up to run something cheaper and / or faster at scale. There's no single best model, only a best model for a budget, a latency target, and a quality floor.

Explore the Aiera Leaderboard >

Aiera Inc. published this content on August 25, 2026, and is solely responsible for the information contained herein. Distributed via Public Technologies (PUBT), unedited and unaltered, on August 25, 2026 at 16:36 UTC. If you believe the information included in the content is inaccurate or outdated and requires editing or removal, please contact us at [email protected]