Frontier Data, Off the Shelf

Training data, evals, and RL environments built by the same experts behind our frontier-lab work. Published gains on public benchmarks. Available today.

35.2% → 47.2%+12.0ppGLM-4.7 on Terminal-Bench 2.0, after post-training on our coding dataset
24.2% → 33.8%+9.6ppQwen3.5-122B-A10B on Toolathlon, after training on our enterprise-agent dataset
4B → 235B-class60×Qwen3-4B matched Qwen3-235B-A22B-Instruct, after training on our instruction-following dataset

Coding Agents

Agents working inside real repositories and real shells.

More datasets in the full catalog

Enterprise Agents

Multi-step work inside the systems companies actually run on.

More datasets in the full catalog

Multimodal Reasoning

Documents, charts, images, interfaces, audio, and other media from real-world work.

More datasets in the full catalog

Expert Reasoning

Advanced reasoning across STEM, dense instructions, long context, multi-turn tasks, and self-consistency.

More datasets in the full catalog

Train first. Pay if it works.

Labs in our Trusted Program can train on the full dataset, evaluate the resulting model, and pay only when it moves the metrics that matter.