What we actually know about frontier data mixes
Fully open, partial, closed — the disclosure gradient, and the numbers you can build on.
Everything here is something a data owner or a lab cannot currently look up: what is actually in frontier training mixes, what a token has sold for, what kills a deal at counsel, and the code we use to measure any of it. Nothing is gated.
What goes into frontier training mixes, and what it costs.
Fully open, partial, closed — the disclosure gradient, and the numbers you can build on.
The mixture is the output of a search procedure, not of taste. Know what the search optimises.
A hundred billion tokens doing disproportionate capability work is the slot you are competing for.
SFT, preference and RLVR mixes, mapped to what each stage consumes and what it pays for.
Which capabilities stopped improving while models grew — and are therefore data-bound.
Disclosed licensing deals normalised per token and per hour. The first honest price sheet.
The playbook a broker's margin depends on you not having.
Price on replacement cost, not volume. The seven scoring dimensions and their weights in full.
The four documents a buyer's counsel asks for on the second call, and where the chain breaks.
Not "we removed PII" but the verification: recall, free-text sweeps, re-identification testing.
What the buyer must publish, therefore what they will demand from you, therefore what to prepare.
Consent, tooling, what to record at each step, and why the failures are the valuable part.
Deal structures with real terms, and the shift from archive sale to committed feed.
Schema, format, and a datasheet. Cleaning is not a favour to the buyer, it is priced.
Our gates, written as war stories rather than as a rubric.
What the collapse literature actually says, and why 2026 layers synthetic on real rather than instead.
No per-record verdicts. Corpus-level estimates only, calibrated per language.
Six parts, real code, one environment carried end to end.
Observation, action, reward, termination — and why LLM agent environments differ from classic RL.
Docker isolation, state reset, determinism, and the seed that makes a task repeatable.
Programmatic checkers, rubric grading and LLM judges — and exactly where each one fails.
Procedural construction, and what breaks when tasks are generated rather than authored.
Hybrid verifiers, and why a passing test is not a solved task.
Packaging, licensing, and what a buyer will ask you about provenance.
The same argument as the writing, shipped as an import.
Deterministic prechecks and the published weights. Runs locally, sends nothing anywhere.
One file, no network, no dependencies. Recomputes the score and checks the delivered sample.
Matched control, real fine-tune, multiple seeds, three diagnostics instead of one number.
Chain of title, licence compatibility and signed manifests, inside the curation stack you already run.
Environments are a traded asset with no standards. Seven dimensions for the ones you sell.