New Frontier data and RL environments, off the shelf

Post-training runs

Published gains on
public benchmarks

Before-and-after results from our post-training runs, with the research behind each one.

Frontier Work

We trained Kimi K2.7 on our Frontier Work data. It improved on GDPVal.

GDPVal
≥90% of rubrics passing
+16.3pp19.6% → 35.9%
GDPVal
100% of rubrics passing
+1.4pp2.7% → 4.1%

General Agents

We trained Qwen3.5-122B-A10B on Surge agent data. The gains transferred to Toolathlon, τ²-Bench, and BFCL-V4.

Toolathlon
pass@1
+9.6pp24.2% → 33.8%
τ²-Bench
pass@1
+5.3pp54.8% → 60.1%
BFCL-V4
pass@1
+3.5pp55.7% → 59.2%

Professional Documents

We trained Kimi K2.7 on GDP.pdf data. It improved across GDP.pdf and GDPVal.

GDP.pdf
pass@1 · holdout
+13.5pp11.0% → 24.5%
GDPVal
≥90% of rubrics passing
+13.1pp19.6% → 32.7%
GDPVal
100% of rubrics passing
+3.2pp2.7% → 5.9%

Agentic Coding

We trained Kimi K2.7 on Surge coding data. It improved on Terminal-Bench 2.1, DeepSWE, and SWE-Bench Pro.

Terminal-Bench 2.1
+15.0pp67.0% → 82.0%
DeepSWE
+12.4pp31.0% → 43.4%
SWE-Bench Pro
+4.7pp60.1% → 64.8%

Instruction Following

We trained Qwen3-4B-Thinking on our ComplexConstraints data. The 4B model nearly matched Qwen3-235B-A22B-Instruct on instruction following.

AdvancedIF
pass@1
+8.4pp28.2% → 36.6%
MultiChallenge
+10.1pp41.1% → 51.2%

STEM Reasoning

We trained GLM4.7 on our Expert STEM data. The gains transferred to HLE, FrontierScience, PubMedQA, and DeepDive.

Expert STEM
pass@1 · held-out
+24.9pp22.7% → 47.6%
HLE
pass@1 · text prompts
+6.8pp26.8% → 33.6%
FrontierScience
pass@1
+9.3pp29.1% → 38.4%
PubMedQA
pass@1
+13.7pp66.7% → 80.4%
DeepDive
pass@1
+9.1pp57.4% → 66.5%

Customer Support Agents

We trained GLM 4.6 on our CoreCraft customer-support environment. The gains transferred to BFCL, τ²-Bench Retail, and Toolathlon.

CoreCraft
task pass rate · held-out
+11.4pp25.4% → 36.8%
BFCL
Parallel · task pass rate
+4.5pp91.0% → 95.5%
BFCL
Simple · task pass rate
+2.0pp91.5% → 93.5%
τ²-Bench
Retail · task pass rate
+7.4pp68.7% → 76.1%
Toolathlon
task pass rate
+6.8pp18.8% → 25.6%

See what the data actually teaches

Many of these runs use datasets available off the shelf. Train on the full dataset first; pay only if it moves the metrics that matter.

Explore off-the-shelf data