Introducing DAYJOB: Healthcare and DAYJOB: Finance
AI models are getting very good at well-specified tasks. Give them a clear objective, detailed instructions, and a clean set of inputs, and frontier models can do remarkable work.
But real work rarely arrives that way. The hard part is often figuring out what the task even is. Real work sounds more like:
“hey @bob is that ready yet?”
And then you have to figure out what that means.
You find the analysis your colleague sent last week. Open the financial model the team has been working from. Notice that the spreadsheet and the latest memo disagree. Decide which source to trust. Update the analysis. Check everything again.
Two days later, you have something your boss can actually use. Can an agent handle that messiness too?
That's the question behind DAYJOB, our new benchmark suite for professional knowledge work. Today we're releasing the first two benchmarks in the family: DAYJOB: Healthcare and DAYJOB: Finance, comprising 100 expert assignments across two of the most economically important areas of professional work.
And today's frontier agents are still a long way from being able to work a 9 to 5.

The strongest system completes only 48.3% of DAYJOB: Healthcare assignments and 20% of DAYJOB: Finance assignments at our 95% completion threshold.
Where GDPval falls short
OpenAI’s GDPval was an important step toward evaluating economically valuable work. But many of its tasks differ from real professional work in two important ways: they often test routine execution rather than professional judgment, and they spell out in detail what the model must do.
1. Many tasks test routine execution, not professional reasoning and judgment
Many GDPval tasks emphasize mechanical execution rather than the core reasoning and judgment associated with the profession being evaluated. They center on document production, formatting, transcription, copy-editing, and other relatively routine forms of work.
One retail-supervisor task, for example, is largely a PDF-generation exercise. The model is given a Word document containing daily duties, and the prompt asks it to turn those duties into a structured task list in a prescribed format. OpenAI’s expert-produced reference deliverable is essentially that formatted checklist.
Almost none of the occupational judgment is left to the model: the content is supplied, the structure and output are set by the prompt. What remains is primarily layout and transcription.
A Private Investigation Supervisor task is similarly centered on editing and formatting a surveillance report rather than conducting investigative analysis.
These tasks may resemble things professionals sometimes do, but that is different from testing the capabilities that make those professionals valuable. Formatting documents, copy-editing, and following a straightforward list of instructions are useful skills. They are not the same as open-ended professional reasoning: figuring out what matters, what to do, and why.
2. The tasks often specify the step-by-step solution in advance
GDPval also frequently tells models not just what outcome to produce, but how to produce it.
Consider one task for a regional grocery-store director. The model is asked to build a promotion-planning form and training deck. But the prompt already specifies most of the solution:
“Make a PowerPoint deck titled ‘Promo Projection Form’ to train Meat Team Leaders on how to use the new form. Keep it to under 8 slides. The deck should explain what this new tool is, why we’re using it, how stores will get the form and where to find it, what sections they’re expected to fill out, and how the process will work. Please include a sample version of the form with some mock data so they can see exactly how it looks when filled out. End with a recap and leave room for discussion or questions.”
The model is not being asked to decide what an effective training deck should contain. Much of that work has already been done in the prompt. Its job is to execute a strategy that has already been specified.
OpenAI's own experiments show why this matters. They created shortened GDPval prompts that removed details about where information lived, how to approach the task, and how the output should be formatted. Model performance dropped: in OpenAI's words, models particularly struggled to “figure out context.”
But figuring out context is a central part of professional work. Real professionals often have to determine what the actual problem is, what information matters, what is missing, what constraints apply, and what a good solution should look like before they can begin producing an answer.
3. Other data and evaluation problems
Our audit found additional issues:
- The input data can be inconsistent or unrealistic. For example, one task asks the model to compare 39 stores even though every store has identical sales data.
- The rubrics are frequently flawed. Only about 20% of the criteria we reviewed were good as-is. We found vague criteria, redundancy, poor atomization, inconsistent standards, and criteria that expected incorrect answers.
- The expert-produced reference deliverables are often flawed. They include hallucinated values, irrelevant information, missing requested sections, and other basic errors.
- Some tasks expire. GDPval includes tasks dependent on live websites, current prices, or time-bound events. Some links in the dataset are already broken.
GDPval partly sidesteps its rubric problems by using human preference judgments between final deliverables. But automated versions replace those experts with LLM judges, which are themselves imperfect graders with biases around style, framing, and model family.
The broader problem is that successful professional work is not just the ability to execute a well-specified task. Benchmarks that supply the context, decompose the problem, prescribe the structure, and then reward polished execution risk measuring the last step of professional work while omitting the messy, real-world reasoning that precedes it.
From benchmark prompts to real assignments
Consider the difference between:
“Open the latest forecast model. Update revenue using the assumptions in the Q3 operating report, revise the downside case, calculate the resulting leverage ratio, and summarize the changes in a one-page memo.”
and:
“Can you update the forecast before tomorrow's meeting?”
The first assumes you’re a junior intern. It tells you what matters, where to look, what analysis to perform, and what to produce. The second is how you talk to your colleagues at a real job.
DAYJOB is built around the second kind of request, leading to three major differences:
- Shorter, more realistic handoffs. DAYJOB prompts are dramatically shorter: 81 words on average in Finance and 49 in Healthcare, compared with 337 words in GDPval. We deliberately leave more of the scoping, interpretation, and planning to the agent.
- Much larger working environments. A DAYJOB task contains an average of 25.7 input files in Finance and 19.8 in Healthcare. GDPval averages just 1.2. The agent has to determine which files matter, connect information across them, and keep the work coherent as it progresses.
- Much longer horizons. A DAYJOB Finance task is estimated to take a human professional 16.6 hours on average; Healthcare averages 13.6 hours. By comparison, the full GDPval set has a median estimated human completion time of 4 hours.

Built by experts on the job
DAYJOB tasks were built by experts with firsthand experience doing the work being evaluated. We asked them to create assignments grounded in the problems, source materials, decisions, and deliverables they encounter in their fields. Every task then went through a three-layer review process for correctness, completeness, and solvability.
The experts behind DAYJOB include people like:
- A senior portfolio manager at BlackRock specializing in institutional asset management, quantitative portfolio construction, and systematic investing, with an MBA from Wharton.
- A finance professional who has structured debt for large-scale infrastructure and renewable energy projects at HSBC, with an MSc in Economics from the London School of Economics.
- A pediatric endocrinologist and epidemiologist focused on diabetes care in children and adolescents, with an MD and MPH from the University of Michigan and a faculty role at a major children’s hospital.
What does a DAYJOB task look like?
The easiest way to understand DAYJOB is to look at where frontier agents fail.
Finance: missing $25.7 million
In one DAYJOB: Finance task, an agent reviewed a fund’s reports before they were sent to a new client. The assignment was simple: find anything that should stop the reports from going out.
One position had been priced manually from a Bloomberg screenshot. The screenshot showed the price in South African cents (“ZAr”), but the value had been entered as rand.
That one unit error made a roughly $260,000 position appear to be worth $26 million.
GPT-5.6 found two real pricing-source errors worth a net $132,000 and correctly recommended holding the reports until they were fixed. But it missed the much larger problem: “No other release blocker was identified in the supplied pack.”
It still missed a $25.7 million overstatement.
Only 2 of 66 runs across 22 model configurations caught the cents-versus-rand error. Both were Claude Opus 5 runs. The successful reviews noticed the unit on the source screenshot, recalculated the position value, and in one case traced the mistake back to an incomplete pricing-system setup.
Healthcare: uncaught dangers
In one DAYJOB: Healthcare task, an urgent-care clinician described a patient’s facial weakness as “textbook Bell’s palsy” and asked the agent to check his laboratory results before discharge.
The chart also showed difficulty speaking and neck pain after moving heavy furniture, and the facial examination was incomplete. Those findings raised concerns that warranted further assessment before discharge.
GPT-5.6 noticed this in its reasoning. It recognized that the combination of facial weakness and neck pain could have a vascular cause. But it decided “I shouldn't question the diagnosis unless necessary.”
A Grok run reviewing the same record did challenge the plan. It connected the incomplete exam and neck pain after physical strain to a possible arterial injury and recommended holding discharge for emergency evaluation and vascular imaging.
These examples capture what makes DAYJOB different. Agents are already very good at calculations, extraction, and other well-specified steps. The harder part is professional judgment: knowing which details matter, what assumptions to make and which to throw out, and which issues are important enough to change a decision.
Frontier agents still have a long way to go
DAYJOB uses task-specific rubrics to measure what the agent completed. The median task has 57.5 criteria in Finance and 47.5 in Healthcare. We count a task as completed when the agent satisfies at least 95% of its criteria; we also report perfect completion.

On DAYJOB: Healthcare, the strongest system achieves 82% average criteria completion and completes 48.3% of assignments at our 95% threshold.
On DAYJOB: Finance, the strongest system achieves 81.7% average criteria completion and completes 20% of assignments at our 95% threshold.
Pour yourself a cup of ambition
Models can already do impressive work when the task is clear: summarize a document, clean up writing, run a calculation, build a slide.
The next milestone is end-to-end professional judgment: figuring out what needs to be done, making judgment calls, and reasoning through the work over days. That’s what DAYJOB measures.
Today we’re launching DAYJOB: Healthcare and DAYJOB: Finance, with more domains to follow.
AI can solve Navier–Stokes. Can it survive a 9 to 5?
Explore DAYJOB: Healthcare and DAYJOB: Finance.







