Benchmarks
This is our ranking of models, measured by their capacity for rigorous reasoning and real-world mastery.
Chartography
Chartography is our benchmark for professional chart understanding. It tests whether frontier models can read the Kaplan-Meier curves, candlestick charts, contour maps, Sankey diagrams, Bode plots, and other specialized graphics that professionals use to make real decisions every day. It evaluates visual perception, domain-aware interpretation, and multi-step graphical reasoning.










































































HANDBOOK.md Agents
Can an agent follow a 100-page company handbook inside an enterprise RL environment?
HANDBOOK.md is a benchmark for long-context agentic instruction following, modeled on how professionals follow corporate policy in their day-to-day work. Each task is a unique RL environment with internal tools and external MCP servers, across five enterprise domains.


































































Antidote: Everyday Edition
Some leaderboards reward the answers that convince you in two seconds; they optimize for confidence, decoration, and flattery. Antidote rewards the answer that leaves you better off a month later; it optimizes for you.
It's our real-world AI leaderboard, graded by doctors, lawyers, and engineers who read every word, check every citation, and run every line of code.
Today's release covers everyday chatbot use, with agentic and enterprise editions coming soon.

































Hemingway-bench
Most AI writing benchmarks reward surface-level signals: elaborate metaphors and prose that looks impressive at a glance. Hemingway-bench rewards writing that is actually good.
Our leaderboard is judged by professional writers who evaluate creative writing, business writing, and everyday writing tasks for taste, originality, coherence, and emotional intelligence.

































ComplexConstraints
A benchmark for the kind of instruction following professional work demands — where constraints depend on each other, fire conditionally, and must be inferred from context.














































GDP.pdf
Can frontier models master the documents that run the world? GDP.pdf is a multimodal and reasoning benchmark that takes real-world prompts and PDFs pulled directly from expert professional workflows.












































EnterpriseBench: CoreCraft Agents
Stop testing models in tiny, self-contained environments. We built CoreCraft, a large-scale startup world, and deployed AI agents to solve real tasks. Our goal: to move agents beyond the cleanliness of the lab and into the chaos of enterprise reality.






Riemann-bench
We evaluate AI models on advanced mathematical problems requiring deep reasoning and novel synthesis. Our benchmark features problems from cutting-edge mathematics, sourced from leading mathematicians – Ivy League professors, PhD IMO medalists, graduate students at the top of their field – in the course of their research.

















































