Frontier Data, Off the Shelf

Training data, evals, and RL environments built by the same experts behind our frontier-lab work. Proven gains on public benchmarks. Trusted Program partners can train first, and pay if it works.
Available today.

24.2% → 33.8%+9.6ppQwen3.5-122B-A10B on Toolathlon,
after training on our enterprise-agent dataset
4B → 235B-class60×Qwen3-4B matched Qwen3-235B-A22B-Instruct, after training on our instruction-following dataset
35.2% → 47.2%+12.0ppGLM-4.7 on Terminal-Bench 2.0,
after training on our coding dataset

Coding Agents

Agents working inside real repositories and real shells.

More datasets in the full catalog

Enterprise Agents

Multi-step work inside the systems companies actually run on.

More datasets in the full catalog

Multimodal Reasoning

Documents, charts, images, interfaces, audio, and other media from real-world work.

More datasets in the full catalog

Expert Reasoning

Advanced reasoning across STEM, dense instructions, long context, multi-turn tasks, and self-consistency.

More datasets in the full catalog

Train first. Pay if it works.

Labs in our Trusted Program can train on the full dataset, evaluate the resulting model, and pay only when it moves the metrics that matter.