We post-trained Kimi K2.7 on 1,465 tasks from our GDP.pdf companion dataset. During training, the model worked on one professional PDF document at a time and used no tools.
As expected, it got much better at GDP.pdf. Strict full-task success (all criteria passed) on the held-out benchmark set rose from 11.0% to 24.5%, moving Kimi K2.7 from 35th to 9th on our public leaderboard.
GDPval tests similar professional reasoning in a different environment. Instead of receiving one document directly, the model has to explore a workspace, decide which sources matter, use tools, perform calculations, and produce a finished artifact. Most tasks contain no PDFs at all. Success on GDPval more than doubled.
The improvements weren't limited to tasks involving PDFs, and the model didn't simply ingest the same information but reason its way to a better final result. It changed how it explored the workspace and gathered evidence before creating deliverables.
On one GDPval task, the base model opened the first weekly timekeeping sheet it found and immediately started building a monthly report. It even noticed the numbers looked suspicious, but kept going.
In contrast, the post-trained model first found all four weekly sheets, combined them, and only then started calculating.
Across the trajectories we analyzed, the trained model looked for more of the relevant sources before acting, even while reading substantially less overall.
Excel spreadsheets weren't part of training. What appears to have transferred was a more general ability to work out what files and information the task depended on, and assemble the complete set before acting.
Think of a junior analyst who learns to check for missing information when reviewing reports. When their next assignment involves a folder of spreadsheets, they apply the same habit. They already know Excel; what transfers is knowing what to look for and check before drawing conclusions.
Transfer from GDP.pdf to GDPval
After post-training Kimi K2.7 on 1,465 GDP.pdf companion tasks, its performance improved on both the holdout set and GDPval:

Both benchmarks test professional reasoning, but the working conditions are different. Instead of handing the model one PDF, GDPval gives it a workspace. The model needs to
- inspect spreadsheets, documents, and other files;
- determine which ones matter;
- use tools;
- run calculations;
- and produce a finished report, spreadsheet, presentation, or PDF.
Across all 220 GDPval tasks, the share scoring at least 90% rose from 19.5% to 32.7%. The stricter full-task pass rate, requiring a 100% score, rose from 2.7% to 5.9%.

PDF tasks improved most, but the gains extended further
One obvious explanation for the transfer is that GDP.pdf taught the model to reason better about PDFs. While GDPval tasks containing PDFs showed larger gains, tasks without PDFs also improved.

At the stricter 100% threshold, the counts are small but point the same way: 1 → 3 of 38 PDF tasks and 5 → 10 of 182 tasks without PDFs.
To understand what changed beyond the scores, we compared how the base and trained models worked through GDPval tasks.
The trained model read less and found more.
The trained model read less, yet found more of what mattered. Training changed what the model looked for and when it started building.
Before training, Kimi K2.7 would often find one plausible source and immediately begin calculating or building the deliverable, leading it to make mistakes and backtrack later on. After training, it was more likely to form the full set of evidence first:
What sources of evidence exist? Which ones matter? Do I have the whole record?
That shows up in the way it used tools:

The trained model accessed more of the relevant sources of information before it began building the final deliverable, while retrieving almost half as much text overall.
It also repeated less of what it had already seen. Its reasoning became less fragmented, with fewer and longer reasoning blocks connecting source discovery to calculations and artifact construction.
This wasn’t brute-force retrieval. Instead of reading more, it got better at assembling the right evidence.

Both GDP.pdf and GDPval require models to assemble evidence
Examples from both benchmarks point to the same underlying shift. The trained model was better at working out what the source information represented, whether it covered the full task, and when missing details or inconsistencies should change the answer. Those skills mattered in both settings: reasoning over a supplied PDF and navigating a workspace with tools to produce a deliverable.
The following pair of tasks illustrates how the same reasoning can help in both environments.
A GDP.pdf task asked the model to calculate the parts required for two complete sets of machine rolls. The manual defined a complete set as six stations with two rolls each. The base model found a single station and treated it as the whole machine. The trained model established the full six-station layout and scaled the parts list from there.
A GDPval workload task showed the same kind of scope error in a very different environment. The model had access to a stakeholder registry, a March budget, and a timekeeping workbook with four weekly sheets. No PDFs were involved.
The base model opened the first weekly sheet and immediately began producing totals for the entire month. It noticed that something looked wrong:
“Everyone is under 40% utilization. That's quite low. Let me verify by checking if maybe the timekeeping data is weekly or something.”
But it never performed that check. Instead, it continued working from the incomplete data.
The trained model first identified the full month’s evidence:
“The timekeeping export has 4 sheets: Mar. 3, Mar. 10, Mar. 17, Mar. 24.”
It combined all four weeks before calculating employee and department workloads and comparing project hours against the budget.

In both tasks, the base model found one relevant piece, treated it as the whole, and prematurely began calculating. The trained model first established the full scope of the task: all six machine stations in GDP.pdf, and all four weeks of timekeeping data in GDPval.
That is the behavior we see transferring across environments.
More matched pairs
Other tasks show similar improvements in assembling and tracking the most relevant information. The base model often worked from part of what the task required: one record instead of the full set, one constraint without propagating its consequences, or some requested elements without carrying all of them into the final output. The trained model was more likely to assemble every part into the whole task.
Keeping the full record
GDP.pdf. A five-year executive-pay table contained six records because two executives served during the same year. The base model kept only one executive per year in its percentage calculations; the trained model included all six.
GDPval. A workstation checklist contained 17 items that needed to be expanded into an action tracker. The base model turned one item into a sample issue and left the remaining rows blank. The trained model carried all 17 items into the tracker, each with action and status fields.
Propagating constraints
GDP.pdf. The base model correctly calculated that a factory production run required 29 hours, then assigned it only five hours in the schedule. The trained model made the schedule consistent with the quantities it planned to produce.
GDPval. An apartment-preparation plan required an additional day for refrigerator installation. The base model acknowledged the delay but left the apartment’s ready date unchanged; the trained model propagated the delay and moved the date back one day.
Maintaining every requirement
GDP.pdf. One task required filtering a long table by both product category and a numeric cutoff. The base model missed qualifying records and included others below the cutoff. The trained model returned all 600 matching records with the correct values.
GDPval. A one-page investor summary had several required elements. The base model fit it onto one page by dropping requested details and cutting off sentences. The trained model satisfied both the content requirements and the one-page constraint.

What transferred?
GDP.pdf did not teach Kimi K2.7 how to inspect files, use tools, and create deliverables. What post-training taught the model is how to approach the evidence needed in professional tasks.
A real-world, professional PDF is not one clean piece of information. The facts needed to solve a GDP.pdf task may be scattered across sections, tables, figures, images, footnotes, and appendices. The model has to work out
- where the relevant information is in the source,
- which details matter,
- how the pieces relate to one another,
- whether it has found enough to act,
- and whether the final answer remains consistent with the evidence.
In that sense, a PDF is a small evidence environment.
GDPval presents the same problem in a different setting. Instead of assembling evidence scattered within one document, the model has to assemble evidence spread across a workspace: deciding which files, sheets, records, and constraints matter before building the deliverable.
Training made Kimi K2.7 less likely to start from the first plausible piece of evidence it found. It was more likely to establish the full scope of the task first and keep that picture in mind as the work progressed.
The key lesson isn’t simply “PDF training makes models better agents.” It’s that training on complex, realistic tasks can teach reasoning skills that carry over to new settings.
Capabilities can outlive the training environment
In an earlier post-training experiment, we trained a model on office work with no coding tasks and still improved on SWE-Bench Pro.
Here, the subject matter remained related, but the way the model worked changed. The training environment was narrow: one professional document, supplied directly, no tools.
The evaluation environment required workspace exploration, tool use, and a finished professional deliverable.
This depended on the model already knowing how to operate in that environment. Training on PDFs did not teach it how to work with new file types or tool interfaces. But it helped it use its existing tools more effectively.
That gets at the question we care about most in post-training: how do we build rich experiences that teach capabilities that generalize across diverse environments? Building intelligence isn’t about climbing individual hills; we want to teach the underlying skills that form a strong foundation.
Before training, Kimi often had enough knowledge and enough tools. The failure was more basic: it would start working before it understood what the job depended on.
A junior analyst does this too. They open the first spreadsheet that looks relevant and start building the deck. A better analyst first asks: what am I missing?
After 1,465 GDP.pdf tasks, Kimi asked that question more often, and that skill survived when the PDF disappeared.
Want to train on the same GDP.pdf companion tasks?
The dataset used in this run is available from Surge.
Subscription confirmed
Appendix: Training run configuration
- Training data: 1,465 GDP.pdf tasks, with roughly ten natural-language grading criteria per task and as many as 30.
- Training method: LoRA in two stages. On-policy self-distillation first pulled the model toward a version of its own response informed by private rubric feedback. Rubric-oriented reinforcement learning then supplied dense rewards against those criteria. GPT-5.4 Mini at medium reasoning served as the judge.
- Training environment: The model received extracted PDF text directly in the prompt, plus page images where non-text content mattered. It did not use tools.
- Evaluation: The final checkpoint was tested on held-out GDP.pdf tasks and GDPval through the Inspect Evals harness. GDPval required the model to inspect available source material, use tools, and produce deliverables such as spreadsheets, reports, presentations, and PDFs. Many tasks contained no PDF inputs.







