Sell2Labs

Selling to Labs

Detection, and why we don't trust it

Our published position

We screen every corpus for undisclosed machine-generated content, because buyers need to know and sellers who disclose deserve not to be undercut by sellers who do not. We also refuse, as a matter of policy, to tell anyone that a particular record was machine-written. Those two positions are consistent, and the reasoning matters more than the conclusion — because a lot of this market is being sold detection products that cannot do what they claim.

What detectors actually measure

Statistical detectors are, in essence, typicality estimators. They ask how predictable a text is under a language model: low perplexity, low variance in sentence structure, a narrow distribution of word choices. Machine-generated text tends to be typical, because generation samples from the high-probability region.

The failure is immediate once you write it down that way. Typicality is a property of writing, not of authorship. Plenty of human writing is highly typical — formal, templated, edited, written to a house style, written by someone drafting carefully in a second language. Plenty of machine writing is atypical, especially when it has been prompted for style or edited afterwards.

The inversion on non-native English

The typicality axis does not merely lose accuracy on non-native English writing. It inverts: the properties that make text read as machine-generated — restricted vocabulary, regular structure, conventional phrasing — are the properties of careful writing in a language that is not your first.

This is the finding that decided our position. A detector applied uniformly across a corpus does not distribute its errors randomly; it concentrates them on a specific population of writers. In a commercial setting, "this record is 89% likely machine-generated" applied to a corpus written largely by non-native English speakers is not a measurement. It is a systematic penalty on a group, laundered through a number.

Any detector we use is therefore calibrated per language and per writer population, with the calibration set documented — and where we cannot calibrate, we report that we cannot, rather than reporting a number we know is biased.

Why per-record verdicts are the wrong product

Even a very good detector fails at the record level for a structural reason. Consider a detector with 95% accuracy on a corpus that is 5% synthetic:

100,000 records, 5% synthetic (5,000), detector at 95%/95%

  true positives      4,750    synthetic, flagged
  false positives     4,750    human-written, flagged
  ----------------------------------------------------
  flagged             9,500    of which 50% are wrong

A record you are told is machine-generated is a coin flip.
And "95% accurate" is a generous assumption for real corpora.

Base rates dominate. The same detector that is nearly useless per record is perfectly serviceable in aggregate: 9,500 flags against an expected 5,000 tells you something real about the corpus, once you have measured the detector's false-positive rate on comparable human text.

What we do instead

  1. Corpus-level estimates with intervals. "Estimated synthetic share 6–14%, at 90% confidence, calibrated against a held-out human corpus from the same domain and language." Never a point estimate, never a per-record label.
  2. Marker screening rather than classification. Assistant preambles, refusal boilerplate, characteristic list structures, template regularity, and near-duplicate generation patterns. These are evidence of process, not of authorship, and they are far more robust than perplexity — a corpus with 3,000 records opening in the same seven ways was generated by something, and that inference does not require a detector.
  3. Per-language calibration, stated. Every screening result names the language and the calibration corpus. An uncalibrated language gets no number.
  4. Disclosure beats detection. A seller who declares "12% of this corpus is model-rephrased, here is the method and the prompt family" scores better on hygiene than one who declares nothing and screens clean. Undisclosed synthetic content is a hygiene failure; disclosed synthetic content is a documented property.
  5. The finding goes to the seller first. If screening suggests undisclosed synthetic content, the seller is told before any buyer is, with the evidence, and gets to respond. Screening output is not published as a verdict on anyone.

What this costs us

It would be commercially convenient to sell a per-record "authenticity score". Buyers ask for one. It would look rigorous on a certificate and it would be entirely spurious, and the first time it labelled a careful human writer as a machine we would deserve everything that followed.

So the position is published rather than buried: no per-record verdicts, corpus-level estimates only, per-language calibration, and disclosure treated as better than a clean screen. If we ever change it, the change will appear here with its reasoning, and certificates will record which policy version they were issued under — the same discipline we apply to the rubric itself.

If you can show that our screening mislabelled a corpus — particularly one written in a language or register we calibrated badly — that is the report we most want. It gets published with the correction.

List a dataset All writing