The Mix Report
The saturation board
A lab buys data for one of two reasons: to defend a capability it has, or to fix one it cannot fix any other way. The second is where the budget is, and it is findable from public numbers. If a capability area has been flat for twelve to eighteen months while model scale, compute and every other input grew, the constraint is not the recipe. It is the data. This is the method we use to find those areas, published so that it can be argued with.
The four columns
| Column | Definition | Why it is in the board |
|---|---|---|
| Ceiling | Best publicly reported score in the area, across all labs, on the evaluations we track for it. | Establishes where the frontier actually is, rather than where any one vendor says it is. |
| Headroom | Distance from the ceiling to the practical maximum — not 100%, but the estimated fraction of items that are answerable given label noise and ambiguity. | An area at 91% against a practical maximum of 93% is finished. Reporting raw distance to 100 manufactures opportunity that is not there. |
| 12-month delta | Change in the ceiling over the trailing year. | The core signal. Compute grew; if the score did not, something else is binding. |
| Cross-lab spread | Gap between the best and median frontier result in the area. | Separates "hard for everyone" from "one lab found something" — which are opposite commercial situations. |
How to read the combinations
The columns are only useful jointly. Four cases, and each implies a different sales motion:
| Delta | Spread | Reading | What it means for a seller |
|---|---|---|---|
| Flat | Narrow | Data-bound. Everyone is stuck in the same place. | The prize case. Every lab has the same hole and none can engineer out of it. Bring evidence of movement and you have several buyers, not one. |
| Flat | Wide | Someone has a private advantage — a corpus, a pipeline, or an eval nobody else fits. | Sell to the laggards, who know exactly what they are behind on. Expect the leader to decline politely. |
| Rising | Narrow | Active competition on a shared recipe. | Fast-moving, but the buyer's problem is being solved without you. Data has to be clearly better than what the recipe is already producing. |
| Rising | Wide | Early area, methods still diverging. | Too early for a licensing deal, right for a partnership where your data shapes the eval. |
The three traps, and how we handle them
A flat line is not automatically a data problem. Three alternative explanations have to be excluded before a capability goes on the board as data-bound, and stating them is what separates a board from a marketing chart.
- Measurement ceiling. Many benchmarks contain items whose gold labels are wrong or ambiguous. A score that has stopped moving may have hit the noise floor rather than the capability limit. We estimate the practical maximum by auditing a sample of items rather than assuming 100%, and an area whose headroom is inside that estimate is marked exhausted, not data-bound.
- Contamination. A benchmark that leaked into training sets reports a score that stopped meaning anything some time ago. Where a public suite has known leakage we say so and downweight it, and where a private variant exists we prefer it.
- Nobody is trying. A capability can be flat because it is commercially uninteresting. Flatness plus no hiring, no published work and no product surface is indifference, not constraint. Hiring is the tiebreaker here: labs advertise for the expertise they are short of before they publish anything about it.
What the board is built from
- Published model cards and technical reports, taken at their own stated evaluation settings.
- Independent leaderboards and reproductions, preferred over vendor self-reports where both exist.
- Our own evaluations of open-weight models where a public number is missing or unreproducible, run with fixed prompts and repeated sampling so the confidence interval is real.
- An audited item sample per evaluation, for the practical-maximum estimate.
Every entry carries its source and its date. Where a number is a vendor self-report that nobody has reproduced, it is marked as such and does not set a ceiling on its own.
Using the board as a seller
- Find your capability area, not your industry. Buyers do not have a "legal" line item; they have long-document reasoning, citation faithfulness, and multi-step procedural execution. Translate.
- Check the delta before you pitch. Pitching into a fast-rising area means competing with the buyer's own roadmap. Pitching into a flat, narrow-spread area means arriving with the one thing they cannot produce internally.
- Bring an evaluation, not an assertion. A held-out split of your own data, published as a benchmark with confidence intervals, converts "we have good data" into "here is a capability nobody scores well on". That artefact is what the board is for, and it is what turns a listing into an inbound.
The uncomfortable corollary: if your area is rising fast and the spread is narrow, the honest answer is that your data is worth less this year than it was last year. We would rather publish a method that can say that than a board that always finds an opportunity.
The numeric board is refreshed quarterly and every cell is sourced. Disputed entries are published with the dispute attached — if you can show a number is stale or a benchmark is contaminated, that correction improves the board.
Read next: What one token costs — the price side of the same question.