New Frontier data and RL environments, off the shelf

Frontier Data
Off the Shelf

Training data, evals, and RL environments built by the same experts behind our frontier-lab work. Proven gains on public benchmarks. Trusted Program partners can train first, and pay if it works.
‍
Available today.

5.0%25.0%+20.0ppKimi K2.7 on SWE-Marathon, after training on our coding dataset
19.6%35.9%+16.3ppKimi K2.7 on GDPVal, after training on our enterprise agents dataset
1.4%12.1%+10.7ppKimi K2.7 on Terminal-Bench 3.0, after training on our coding dataset

Coding Agents

Agents working inside real repositories and real shells.

More datasets in the full catalog

Enterprise Agents

Multi-step work inside the systems companies actually run on.

More datasets in the full catalog

Multimodal Reasoning

Documents, charts, images, interfaces, audio, and other media from real-world work.

More datasets in the full catalog

Expert Reasoning

Advanced reasoning across STEM, dense instructions, long context, multi-turn tasks, and self-consistency.

More datasets in the full catalog

Train first. Pay if it works.

Labs in our Trusted Program can train on the full dataset, evaluate the resulting model, and pay only when it moves the metrics that matter.