The Mix Report
What one token costs
There is no price sheet in this market, which suits everyone selling intermediation. There are, however, roughly a dozen deals whose value has been disclosed or credibly reported, and a labour market for expert-produced data with published rates. Normalise both onto the same axis and the structure of the market becomes obvious — and so does which side of it you are on.
The disclosed deals
| Seller | Reported value | Shape |
|---|---|---|
| News Corp | ~$250M over five years | Archive plus ongoing feed, multi-title |
| ~$203M aggregate across agreements | Continuous API access to user-generated content | |
| Wiley | $40M+ across two agreements | Academic book and journal content |
| Axel Springer | ~$13M over three years | News archive plus feed |
| Financial Times | ~$5–10M per year | Archive plus feed, single title |
Values are as reported publicly; several are estimates from filings or press coverage rather than confirmed contract values, and none of these counterparties has confirmed a per-token rate. Treat the table as the shape of the market, not as a rate card.
Normalising, with the assumptions on the table
Nobody discloses corpus size alongside price, so per-token figures require an estimate of the archive. The honest way to do this is to state the estimate and show the sensitivity rather than to publish a single confident number.
Method
price per 1k tokens = annualised deal value
÷ (estimated corpus tokens / 1,000)
Worked, for a mid-size national newspaper archive
annualised value $8M/yr
articles in archive ~3M (say 40 years at ~75k/yr)
tokens per article ~800
corpus ~2.4B tokens
→ $8M / 2.4M thousand-token units ≈ $3.3 per 1k tokens ... first year
But the archive is licensed for a term, and the buyer amortises it across
the whole term while the feed adds new tokens each year:
over 5 years, ~2.4B archive + ~0.3B new ≈ 2.7B tokens for ~$40M
→ ≈ $15 per million tokens = $0.015 per 1k tokens
Change the archive estimate by 2× and the number moves by 2×. That is the honest error bar. What survives the uncertainty is the order of magnitude: published text archives clear somewhere in the range of cents to tens of cents per thousand tokens, and the largest deals are large because the corpora are large, not because the rate is high.
The other axis: data produced to order
Expert data is not priced per token at all. It is priced per hour, at rates that have been publicly reported around $85 per hour averaged across domains, rising to roughly $200 per hour for physicians and comparable specialists. To compare, convert:
An expert producing worked, reasoned output generates on the order of
600–1,200 tokens of usable content per hour once review is included.
$85/hr ÷ ~900 tokens/hr ≈ $94 per 1k tokens
$200/hr ÷ ~900 tokens/hr ≈ $222 per 1k tokens
| Category | Order of magnitude, per 1k tokens |
|---|---|
| Open web text, post-filtering | Effectively free — the cost is compute, not licence |
| Licensed published archives | $0.01 – $0.50 |
| Specialist archives with restricted supply | $1 – $20 (thinly evidenced; few disclosures) |
| Expert-produced demonstrations and judgements | $50 – $250 |
Three to four orders of magnitude separate the top and bottom rows. That spread is the single most important number in this report, and it explains almost every behaviour in the market: why labs will spend nine figures on archives and still complain about data scarcity; why annotation firms command the valuations they do; and why a data owner sitting on a specialist archive is holding something priced between two very different worlds.
Four things the per-token frame gets wrong
- It ignores rights quality. Two corpora at the same rate are not the same asset if one comes with clean chain of title and indemnity and the other does not. A large part of what the big deals bought was not tokens but defensibility — see chain of title.
- It ignores exclusivity. A non-exclusive licence to a corpus that five labs will also license is worth a fraction of the same corpus sold once. Structure moves price more than volume does; that is its own post.
- It ignores the feed. Most large deals are archive plus ongoing access, and buyers increasingly value the feed above the archive because a static corpus ages. A committed refresh cadence is the difference between a sale and a subscription.
- It ignores where the tokens land. A token in the pretraining mix and a token in the post-training set are different goods with different economics. Pricing per token averages across a distinction that is worth two orders of magnitude — see post-training.
How to use this if you are selling
- Work out which row of that table you are in. Most sellers assume row three and are in row two.
- If you are in row two, your leverage is volume, rights quality and refresh — not uniqueness.
- If you can move to row four by producing rather than only archiving — running a judgement programme, or a capture programme over work already being done — the rate difference is large enough to justify building the capability.
- Anchor negotiations on replacement cost and use per-token figures only to sanity-check the result. Per-token is a comparison tool, not a valuation method.
Corrections wanted, especially on corpus-size estimates: every per-token figure here is a division by a number we estimated. If you know the real denominator for a deal in the table, that correction gets published with attribution.