Selling to Labs
The five ways a deal dies
Our intake gates exist because of specific failures, and a rubric is a poor way to explain them. These are the five, in roughly the order they occur, told as what actually happened. Details are changed and cases are composites, but nothing here is hypothetical — each one cost somebody a deal that should have closed.
One · The contractor who owned the work
A services firm had eleven years of annotated case files. Genuinely scarce, genuinely useful, clean schema, a buyer already interested. Diligence asked the routine question about who produced the annotations. The answer was: a rotating pool of specialist contractors, engaged on a one-page purchase order that said nothing about intellectual property.
Under the applicable law the contractors owned their output. The firm had a licence to use it internally by implication and nothing more. Reaching 40-odd contractors, several retired and two deceased, for a retroactive assignment was not a project anybody would fund against an uncertain sale. The deal ended — not with a no, but with three months of silence and then a polite decline.
The gate: we ask for the assignment clause and evidence of coverage before anything else, because it is the failure with no remedy. Chain of title is the checklist.
Two · The duplicate rate the seller had never measured
A listing claimed 4.2 million support conversations. Genuine records, real customers, correct rights position. The buyer's first act was to run deduplication, which is a twenty-minute job. 61% of the corpus was near-duplicate: templated confirmations, automated replies, and the same three flows repeated at scale.
The effective corpus was 1.6 million conversations, which is still a real asset. But the seller had opened at a price anchored on 4.2 million and had said "unique conversations" in writing. The conversation stopped being about the data and became about whether the seller knew their own product. It closed eventually, at a third of the opening, after two extra months.
The gate: effective volume after near-duplicate discounting is computed at intake and shown on the listing. It costs a seller nothing to know this first and a great deal to learn it from the buyer.
Three · The benchmark that was in the training data
An education company's question bank showed a large, clean lift on a public reasoning benchmark in the buyer's pilot. Everybody was pleased for about a week. Then an engineer ran an n-gram overlap check and found that a subset of the corpus was derived from the same public item pool the benchmark was drawn from. The lift was partly memorisation of the test.
The corpus was still useful, and a decontaminated version would have shown a smaller, real effect. But the number that had generated the enthusiasm was gone, and the pilot had to be rerun on the buyer's budget. Nobody accused anyone of anything; the deal simply lost the momentum that had carried it.
The gate: contamination screening against public suites before any capability claim is published. A lift that turns out to be leakage is worse than no lift at all, because it retroactively taints every other number you have quoted.
Four · The sample that was not the corpus
A seller supplied a 5,000-record sample. Excellent: rich narration, consistent labels, hard cases well represented. Terms were agreed. The full delivery arrived and the narration field was empty on 78% of records — the sample had been drawn, understandably, from the period when a particular team had been diligent about filling it in.
The seller had not intended to mislead. They had sent their best work, which is what people do. But the buyer had priced the corpus on the sample, and what arrived was a different asset. It went to a renegotiation that neither side enjoyed, and the relationship did not survive it.
The gate: samples are drawn by us, deterministically, from the whole corpus, and the certificate records the hash of exactly what was assayed. A buyer can check the delivery against it with s2l-verify. This protects honest sellers more than it catches dishonest ones.
Five · The terms of service that did not stretch
A platform with a decade of user-generated content, a strong rights story on paper, and a buyer who very much wanted it. The terms of service granted a licence to use content "to operate and improve the service". The corpus was largely created before 2021, under an earlier version of the terms with narrower language.
The buyer's counsel would not accept "improve the service" as covering third-party model training, in a context where the buyer has to publish a description of its training sources. The seller offered indemnity. Counsel's answer was that indemnity covers money, not the obligation to describe the corpus accurately in public. The deal died on that distinction, and it is the distinction that has changed this market most in the last two years.
The gate: we ask which version of the terms was in force when the records were created, and whether opt-outs were applied to the export. Article 53 explains why the answer has stopped being negotiable.
What the five have in common
| Failure | Cost to check first | Cost of being found out |
|---|---|---|
| Contractor IP gap | A day in the filing system | The deal, with no remedy |
| Duplicate mass | Twenty minutes of compute | Two-thirds of the price |
| Benchmark contamination | An hour of n-gram overlap | Every number you have quoted |
| Unrepresentative sample | One honest random draw | The relationship |
| Terms that do not cover training | Reading your own terms with a date in hand | The deal, late and expensively |
Every one is cheap to find and expensive to be told. None of them is about whether the data is good. That is the part sellers find hardest to accept: the quality of your corpus is rarely what decides the outcome, because by the time quality is being discussed you have already passed the gates that kill most deals.
Read next: Synthetic didn't replace you — the argument sellers hear most often, and what the evidence actually supports.