Selling to Labs
What your data is actually worth to an AI lab
Almost every first conversation we have starts with a volume number. Rows, hours, tokens, gigabytes. It is the wrong opening, because volume is the one property a buyer can obtain elsewhere. The number that governs the deal is replacement cost: what it would cost the lab to have qualified people produce the equivalent from scratch.
Why volume is not the anchor
A frontier pretraining mix is assembled from a pool measured in trillions of tokens, of which a substantial fraction is discarded during selection (see what we know about frontier mixes). Against that denominator, no commercial dataset is large. If your pitch is scale, you have entered a comparison you cannot win and priced yourself against free web text.
Replacement cost inverts the question. Instead of "how much data is this?", ask: if this corpus did not exist, what would the buyer have to do? For most valuable datasets the answer involves people with credentials, doing something over a period of time, under conditions that cannot be recreated on demand. That is the quantity being sold. The rows are the receipt.
The arithmetic
Reported market rates for expert annotation work sit around $85 per hour averaged across domains, rising to roughly $200 per hour for physicians and comparable specialists. Those numbers are the buyer's build option, and therefore your ceiling and your floor at once.
Replacement cost = hours of qualified labour embodied
× the rate that grade of labour commands
× a recreatability factor
Recreatability 1.0 an annotation shop could reproduce this next quarter
1.5 requires rare credentials or access, but is buildable
3.0+ depends on a longitudinal record or an institutional
position that cannot be bought at any price
Worked example. A clinic holds 9,000 triage conversations with the clinician's reasoning recorded at each turn and the eventual outcome attached. Each conversation embodies roughly 20 minutes of physician time — 3,000 hours. At $200/hour that is $600,000 of build cost before any recreatability adjustment. The outcome linkage cannot be produced by an annotation vendor at all, because it required following the patient; call it 2×. Replacement cost lands near $1.2M.
That is not the price. It is the number the price is negotiated against, and it is a far better opening than "9,000 rows". Everything in the rubric below is a discount or a premium applied to it.
The seven dimensions
This is the scoring rubric we run. The weights are published because sellers optimising against them improves the asset — a corpus with better-documented rights and lower duplicate mass is genuinely worth more, not just better-scoring. (The judge prompts behind the qualitative dimensions stay closed, for the opposite reason: optimising against a prompt improves only the sample. The assay post explains where we draw that line.)
| # | Dimension | Weight | What it measures |
|---|---|---|---|
| 1 | Rights & provenance | 24 | Chain of title from origin to you, consent basis, licence compatibility with model training, rights reservations honoured, and whether the documentation would survive a buyer's counsel. |
| 2 | Capability lift | 20 | Measured movement on a held-out evaluation when the data is added to a mix, against a matched control on the same token budget. Not a vibe; a number with a confidence interval. |
| 3 | Scarcity | 16 | Whether an equivalent exists publicly or can be commissioned. This is the replacement-cost multiplier expressed as a score. |
| 4 | Fidelity | 14 | Label and annotation correctness, internal consistency, expert agreement on an audited slice, and whether errors are random or systematic. |
| 5 | Coverage | 12 | Effective volume after near-duplicate discounting, spread across the claimed domain, and representation of the hard tail rather than only the modal case. |
| 6 | Delivery readiness | 8 | Schema stability, format, a datasheet covering collection method and known gaps, and a committed refresh cadence if the asset is a feed rather than an archive. |
| 7 | Hygiene | 6 | Benchmark contamination, undisclosed synthetic content, PII residue, and encoding damage. |
Two features of that table are the whole argument.
Rights is the largest single weight, at 24. Not because provenance is more interesting than capability, but because it is the dimension that fails absolutely. A dataset scoring poorly on coverage is worth less; a dataset with an unresolvable rights position is worth nothing, at any price, to a buyer who has to publish a training-data summary. In our own pipeline it is the most common cause of a deal ending after the second call.
Capability lift is only 20. Sellers expect it to be everything, and buyers behave as if it were, right up until counsel gets involved. A large measured lift on an asset nobody can license is a very expensive way to learn about chain of title.
How the score is used
The composite is a weighted sum on a 0–100 scale, but it is not a price and it is not a ranking against other sellers. It does three narrower things: it sets which dimensions get remediated before a listing goes out, it decides whether a buyer's gates are passed at all, and it gives both sides a shared vocabulary when the negotiation gets to specifics. A certificate carries the composite, the seven component scores, the rubric version, and our relationship to the seller — all recomputable by the counterparty with s2l-verify.
What moves the number most, per unit of effort
- Document the rights you already have. Most sellers score badly here not because their position is bad but because it is undocumented. Finding the source agreements and the contractor IP assignment is a week of work against the heaviest weight in the rubric.
- Deduplicate before you measure. Near-duplicate mass inflates volume and depresses effective coverage simultaneously. It is also the easiest thing for a buyer to check on day one, which makes an inflated volume claim an unforced credibility loss.
- Keep the failures. Corpora are routinely cleaned of the cases where the process went wrong and was recovered. For agentic and procedural data those are the highest-value records in the set, and pruning them is the single most common act of value destruction we see.
- Write the datasheet. Collection method, time span, inclusion and exclusion criteria, known gaps, and what the data does not support. Stating a limitation costs less than having a buyer discover it.
What does not move it
Growing the corpus by adding more of the same thing; converting to a fancier format; generating synthetic expansions without disclosing them (which is detected, and moves the hygiene score down rather than volume up); or obtaining a valuation letter from someone who has not measured capability lift. There is no shortage of parties willing to sell data owners a number. Ask any of them what control they ran it against.
Read next: Chain of title, in practice — the documents a buyer's counsel asks for on the second call.