Sell2Labs

The Mix Report

What one token costs

Edition 1Every figure sourced or marked as an estimate

There is no price sheet in this market, which suits everyone selling intermediation. There are, however, roughly a dozen deals whose value has been disclosed or credibly reported, and a labour market for expert-produced data with published rates. Normalise both onto the same axis and the structure of the market becomes obvious — and so does which side of it you are on.

The disclosed deals

SellerReported valueShape
News Corp~$250M over five yearsArchive plus ongoing feed, multi-title
Reddit~$203M aggregate across agreementsContinuous API access to user-generated content
Wiley$40M+ across two agreementsAcademic book and journal content
Axel Springer~$13M over three yearsNews archive plus feed
Financial Times~$5–10M per yearArchive plus feed, single title

Values are as reported publicly; several are estimates from filings or press coverage rather than confirmed contract values, and none of these counterparties has confirmed a per-token rate. Treat the table as the shape of the market, not as a rate card.

Normalising, with the assumptions on the table

Nobody discloses corpus size alongside price, so per-token figures require an estimate of the archive. The honest way to do this is to state the estimate and show the sensitivity rather than to publish a single confident number.

Method
  price per 1k tokens  =  annualised deal value
                          ÷ (estimated corpus tokens / 1,000)

Worked, for a mid-size national newspaper archive
  annualised value        $8M/yr
  articles in archive     ~3M  (say 40 years at ~75k/yr)
  tokens per article      ~800
  corpus                  ~2.4B tokens
  →  $8M / 2.4M thousand-token units  ≈  $3.3 per 1k tokens ... first year

  But the archive is licensed for a term, and the buyer amortises it across
  the whole term while the feed adds new tokens each year:
  over 5 years, ~2.4B archive + ~0.3B new  ≈  2.7B tokens for ~$40M
  →  ≈ $15 per million tokens  =  $0.015 per 1k tokens

Change the archive estimate by 2× and the number moves by 2×. That is the honest error bar. What survives the uncertainty is the order of magnitude: published text archives clear somewhere in the range of cents to tens of cents per thousand tokens, and the largest deals are large because the corpora are large, not because the rate is high.

The other axis: data produced to order

Expert data is not priced per token at all. It is priced per hour, at rates that have been publicly reported around $85 per hour averaged across domains, rising to roughly $200 per hour for physicians and comparable specialists. To compare, convert:

An expert producing worked, reasoned output generates on the order of
600–1,200 tokens of usable content per hour once review is included.

  $85/hr   ÷  ~900 tokens/hr   ≈  $94 per 1k tokens
  $200/hr  ÷  ~900 tokens/hr   ≈  $222 per 1k tokens
CategoryOrder of magnitude, per 1k tokens
Open web text, post-filteringEffectively free — the cost is compute, not licence
Licensed published archives$0.01 – $0.50
Specialist archives with restricted supply$1 – $20 (thinly evidenced; few disclosures)
Expert-produced demonstrations and judgements$50 – $250

Three to four orders of magnitude separate the top and bottom rows. That spread is the single most important number in this report, and it explains almost every behaviour in the market: why labs will spend nine figures on archives and still complain about data scarcity; why annotation firms command the valuations they do; and why a data owner sitting on a specialist archive is holding something priced between two very different worlds.

Four things the per-token frame gets wrong

How to use this if you are selling

  1. Work out which row of that table you are in. Most sellers assume row three and are in row two.
  2. If you are in row two, your leverage is volume, rights quality and refresh — not uniqueness.
  3. If you can move to row four by producing rather than only archiving — running a judgement programme, or a capture programme over work already being done — the rate difference is large enough to justify building the capability.
  4. Anchor negotiations on replacement cost and use per-token figures only to sanity-check the result. Per-token is a comparison tool, not a valuation method.

Corrections wanted, especially on corpus-size estimates: every per-token figure here is a division by a number we estimated. If you know the real denominator for a deal in the table, that correction gets published with attribution.

List a dataset All writing