Sell2Labs

Selling to Labs

Synthetic didn't replace you

Edition 1

Every data owner hears it, usually from someone with no stake in the answer: the labs can generate their own data now, so yours is worth less. It is a reasonable-sounding claim built on a category error. Synthetic data is a transformation applied to something a model already knows or can verify. It expands, rephrases, and searches — it does not observe. The question is not whether synthetic works. It is what it works on.

What synthetic data actually is

KindWhat it doesWhat it needs to already have
Rephrasing / augmentation Restates existing documents in cleaner or more instructional form The source documents
Distillation Transfers a stronger model's behaviour into a weaker one A stronger model that already has the capability
Self-play with a verifier Generates attempts and keeps the ones a checker accepts A domain where correctness is programmatically checkable
Free generation Produces plausible text in a domain with no grounding Nothing — and this is the kind that degrades models

Three of those four require real data or a real verifier as an input. Only the fourth is the thing sellers are told will replace them, and it is the one that does not work.

What the collapse literature supports

The findings are more nuanced than either camp quotes them as being, and both numbers matter to a seller:

Read together: synthetic data is safe in proportion to how tightly it is anchored to something real, and dangerous in proportion to how free it is. Which makes real data the input that governs how much synthetic can be used at all.

The 2026 pattern

What the fully-open recipes actually do — and what the closed labs' published work implies they do — is layer synthetic on top of licensed and curated real data, at every stage:

  1. Real corpus in. Filtered, deduplicated, quality-classified.
  2. Synthetic rephrasing of that corpus into denser, more instructional form. The facts come from the real documents; the model supplies the phrasing.
  3. Synthetic expansion around real seed examples — variations on real cases, not invented ones.
  4. Verifier-filtered generation for domains where a checker exists, keeping only what passes.
  5. Real held-out data for evaluation, because a model evaluated on its own output measures nothing.

At every one of those steps, more and better real data raises the ceiling of what synthetic can produce. This is why demand for licensed data went up over the same period synthetic techniques matured. The multiplier got better, and a better multiplier makes the multiplicand more valuable, not less.

Where synthetic genuinely does replace you

Being honest about this is what makes the rest credible. If your corpus is any of the following, synthetic generation is a real substitute and your price should reflect it:

What is not substitutable: observation of the world (what actually happened, to whom, with what outcome), expert judgement on genuinely contested cases, the real distribution of requests a professional receives, failures and recoveries, and anything whose correctness depends on facts the model cannot access.

How to answer the objection in a meeting

Not by arguing about collapse papers. By locating your corpus on that split:

"Everything in this corpus is an observation with an outcome attached. A model can generate a plausible triage conversation. It cannot generate what the patient actually turned out to have — and that field is why the corpus exists. Here is what happens to the eval when you drop it."

Then show the measurement. A contribution test against a matched control settles this argument in a way no amount of literature citation will; that is what marginal is for, and if your data really is substitutable, the test will say so — which is worth knowing before a buyer's pilot says it for you.

A related consequence: if you have used synthetic augmentation on your own corpus, disclose the share and the method. Buyers screen for it, undisclosed synthetic content is treated as a hygiene failure, and the detection question is its own post.

Measure the contribution All writing