Post-training runs
Published gains on
public benchmarks
Before-and-after results from our post-training runs, with the research behind each one.
Frontier Work
We trained Kimi K2.7 on our Frontier Work data. It improved on GDPVal.
General Agents
We trained Qwen3.5-122B-A10B on Surge agent data. The gains transferred to Toolathlon, τ²-Bench, and BFCL-V4.
Professional Documents
We trained Kimi K2.7 on GDP.pdf data. It improved across GDP.pdf and GDPVal.
Agentic Coding
We trained Kimi K2.7 on Surge coding data. It improved on Terminal-Bench 2.1, DeepSWE, and SWE-Bench Pro.
Instruction Following
We trained Qwen3-4B-Thinking on our ComplexConstraints data. The 4B model nearly matched Qwen3-235B-A22B-Instruct on instruction following.
STEM Reasoning
We trained GLM4.7 on our Expert STEM data. The gains transferred to HLE, FrontierScience, PubMedQA, and DeepDive.
Customer Support Agents
We trained GLM 4.6 on our CoreCraft customer-support environment. The gains transferred to BFCL, τ²-Bench Retail, and Toolathlon.
See what the data actually teaches
Many of these runs use datasets available off the shelf. Train on the full dataset first; pay only if it moves the metrics that matter.
Explore off-the-shelf data