Sell2Labs

How to sell data to AI labs

Everything here is something a data owner or a lab cannot currently look up: what is actually in frontier training mixes, what a token has sold for, what kills a deal at counsel, and the code we use to measure any of it. Nothing is gated.

The Mix Report

What goes into frontier training mixes, and what it costs.

What we actually know about frontier data mixes

Fully open, partial, closed — the disclosure gradient, and the numbers you can build on.

How a mix actually gets chosen

The mixture is the output of a search procedure, not of taste. Know what the search optimises.

Midtraining is where licensed data lands

A hundred billion tokens doing disproportionate capability work is the slot you are competing for.

Post-training is where the money is

SFT, preference and RLVR mixes, mapped to what each stage consumes and what it pays for.

The saturation board

Which capabilities stopped improving while models grew — and are therefore data-bound.

What one token costs

Disclosed licensing deals normalised per token and per hour. The first honest price sheet.

Selling to Labs

The playbook a broker's margin depends on you not having.

What your data is actually worth to an AI lab

Price on replacement cost, not volume. The seven scoring dimensions and their weights in full.

Chain of title, in practice

The four documents a buyer's counsel asks for on the second call, and where the chain breaks.

De-identification that survives counsel

Not "we removed PII" but the verification: recall, free-text sweeps, re-identification testing.

Article 53 in one page for data owners

What the buyer must publish, therefore what they will demand from you, therefore what to prepare.

How to run a trajectory capture programme

Consent, tooling, what to record at each step, and why the failures are the valuable part.

Non-exclusive, exclusive, first-look

Deal structures with real terms, and the shift from archive sale to committed feed.

Ship it training-ready

Schema, format, and a datasheet. Cleaning is not a favour to the buyer, it is priced.

The five ways a deal dies

Our gates, written as war stories rather than as a rubric.

Synthetic didn't replace you

What the collapse literature actually says, and why 2026 layers synthetic on real rather than instead.

Detection, and why we don't trust it

No per-record verdicts. Corpus-level estimates only, calibrated per language.

Build an RL Gym

Six parts, real code, one environment carried end to end.

What an environment actually is

Observation, action, reward, termination — and why LLM agent environments differ from classic RL.

Wrapping a real application

Docker isolation, state reset, determinism, and the seed that makes a task repeatable.

Reward design

Programmatic checkers, rubric grading and LLM judges — and exactly where each one fails.

Task generation at scale

Procedural construction, and what breaks when tasks are generated rather than authored.

Verification

Hybrid verifiers, and why a passing test is not a solved task.

Publishing an environment

Packaging, licensing, and what a buyer will ask you about provenance.

Tools & Code

The same argument as the writing, shipped as an import.

assay — the rubric, as a command you can run

Deterministic prechecks and the published weights. Runs locally, sends nothing anywhere.

s2l-verify — check our certificate without trusting us

One file, no network, no dependencies. Recomputes the score and checks the delivered sample.

marginal — what would this dataset add?

Matched control, real fine-tune, multiple seeds, three diagnostics instead of one number.

provenance-ops — rights operators for existing pipelines

Chain of title, licence compatibility and signed manifests, inside the curation stack you already run.

env-assay — grading RL environments

Environments are a traded asset with no standards. Seven dimensions for the ones you sell.

List a dataset Post an RFP