Frontier Data, Off the Shelf
Training data, evals, and RL environments built by the same experts behind our frontier-lab work. Published gains on public benchmarks. Available today.
Coding Agents
Agents working inside real repositories and real shells.
Repository-Level Software Engineering
Tasks set inside complete codebases, covering feature development, bug fixes, and production refactors. Compatible with SWE-bench out of the box.
Terminal-Use Agents
Command-line tasks in Linux containers, graded by hidden verifiers. An agent has to inspect the environment and produce something that works: a patched library, a faster training run.
Enterprise Agents
Multi-step work inside the systems companies actually run on.
Toolathlon
RL environments for long-horizon tool use across office documents, enterprise workflows, and live external services.
GDPval-Verified
Long-horizon tasks modeled on real professional work across 30 domains, each paired with a golden response from a domain expert.
Multimodal Reasoning
Documents, charts, images, interfaces, audio, and other media from real-world work.
Enterprise Document Reasoning
Expert-authored prompts over the PDFs that run the economy: filings, invoices, schematics, medical records.
Professional Chart Understanding
Chart-reasoning tasks across healthcare, finance, manufacturing, and STEM, covering survival curves, candlestick charts, contour maps, Bode plots, and other specialized graphics.
Expert Reasoning
Advanced reasoning across STEM, dense instructions, long context, multi-turn tasks, and self-consistency.
Complex Instruction Following
Prompts and rubrics for following dense, multi-part instructions.
Advanced STEM Reasoning
Complex prompts across science, math, engineering, finance, and the social sciences, each with a verified answer and an expert-written explanation.
Train first. Pay if it works.
Labs in our Trusted Program can train on the full dataset, evaluate the resulting model, and pay only when it moves the metrics that matter.