The Research Org Got a Second Workforce
OpenAI says its research organization now uses 3.1 days of agent labor for every day of human labor. The company published the number on Sunday. It is specific enough to sound like an answer. It is actually the beginning of a much harder accounting problem.
OpenAI calls the system it has built an automated research intern. Its definition is narrower than the name suggests: a system that performs well-defined research tasks under human direction, including work that would take a skilled researcher a few days. OpenAI says it reached that milestone in September, on schedule with a goal announced last fall.
By mid-August, according to the company, the median researcher was using more than six hundred dollars of coding-agent inference per day at API prices. The ninetieth-percentile user was consuming more than seven thousand dollars' worth of tokens per day. Researchers increasingly run several agents at once, including subagents launched by the systems they start directly.
Then comes the figure likely to survive every presentation deck: 3.1 agent-workdays for each human workday. OpenAI calculated that from total agent runtime using an eight-hour workday.
Runtime is not labor output. Four agents running beside one researcher can produce useful code, failed experiments, duplicate attempts, abandoned branches, or some mixture of all four. The clock records each one without becoming embarrassed.
OpenAI does not pretend otherwise. Its post says the measurements are preliminary. Code volume and experiment counts are comparatively easy to measure but hard to interpret. The company's available compute has also grown significantly since 2025. More experiments can mean faster research. They can also mean more ways to discover that an idea does not work.
The company's own task data shows the supervision burden. OpenAI says coding agents still need significant human steering, especially on harder work. During the six months it examined, more than half of successful tasks estimated at four to eight hours involved at least one human intervention. The chart excludes tasks whose outcomes were uncertain.
So the useful question is not whether the agents ran for 3.1 times as many hours as the humans. They did, by OpenAI's count. The question is what kind of work those hours bought.
What the agents do
OpenAI classified agent activity using a framework developed by Epoch AI. It divides frontier research into six phases: deciding what to work on, designing research and engineering specifications, building code and datasets, running training and evaluations, analyzing results, and communicating findings and decisions.
Agent use increased across all six categories between January and August. But high-level planning remained a minimal share of agent output. OpenAI says people still set priorities, decide which results deserve pursuit, and choose whether to scale, pause, or deploy a system.
The clearest operational change may be less glamorous. Researchers use coding agents to troubleshoot internal infrastructure. OpenAI says several technical-support teams saw fewer researchers at office hours, and one stopped holding them so its staff could work on other improvements. That is real organizational change: a support queue moved from people answering recurring questions to agents handling more of the first-line work.
It is still a company describing itself. OpenAI published the definitions, the caveats, and some methods, which is better than publishing a victory lap with the denominator removed. Outsiders cannot inspect the internal task mix, rerun the classifications, or connect the runtime figure to accepted research results.
An inspectable comparison
Epoch AI and Proximal offer a useful contrast. Their new Frontier SWE version two benchmark contains thirty-four difficult software-engineering and AI-research tasks drawn from technical domains. Each model gets five attempts per task and as much as twenty hours for each attempt. The public results include mean, best and worst scores, cost, wall-clock time, and traces.
The current leaderboard is not a measure of OpenAI's internal research organization, and it does not audit the 3.1 figure. It shows what inspectable agent-work accounting can look like. A task is defined. Several attempts are visible. Completion receives a score. Cost, time, variance, and traces travel with the result.
The spread is wide. In the September 7 leaderboard snapshot, the leading system had a mean score of 56.3 percent. GPT-5.6 scored 32.2 percent. Those are benchmark scores, not percentages of human researchers replaced. They are also a useful warning against turning hours of machine activity into human-equivalent work by arithmetic.
Epoch's broader research-task framework makes the same problem explicit. It lists more than sixty AI-research tasks and grades automation from zero to five. At level two, AI assists while the human still drives and reviews the work. At level four, AI leads but a human supervises, corrects course, and approves the output. Level five means end-to-end work with little or no human involvement.
An agent-workday does not tell you which level occurred. It does not tell you whether the run succeeded, whether a person repaired it, whether three agents tried the same approach, or whether the output changed a research decision.
Different columns, different results
Older evidence shows why that distinction is not academic. Epoch examined forty-one core contributors to OpenAI's public Codex repository. In the second quarter of 2026, eight percent of contributor-days included merged work that an ensemble of models estimated would take an experienced unaided engineer more than twenty-four hours, up from two percent a year earlier. Epoch calls that estimate an upper bound on time saved. Without AI, engineers might build differently, and more complicated code is not necessarily more valuable code.
A 2025 randomized trial from METR found the opposite result in a different setting. Sixteen experienced open-source developers completed 246 tasks in mature projects. They expected AI tools to make them faster and later believed the tools had made them faster. Measured completion time increased by nineteen percent. The tools had slowed them down.
Neither study settles what is happening inside OpenAI in 2026. Together they show why adoption, runtime, output volume, perceived speed, and completed useful work belong in different columns.
What the ledger needs
Here's what I think. Agent-workday is a useful measure of how deeply a lab has reorganized itself around coding agents. It is a poor substitute for productivity. If a manager claimed output had tripled because three batch processes ran beside every employee, finance would eventually ask what shipped. OpenAI's disclosure proves that the second workforce is operational. Proving what it contributes requires a ledger that records the task, completion test, retries, interventions, accepted output, cost, and the decision that changed because of it.
The phrase research intern is honest only if we keep the supervision in the definition. This intern can be copied, run overnight, and assigned in parallel. It also works inside boundaries chosen by people, on priorities chosen by people, with results judged by people. The copied intern changes the research loop without becoming the research organization.
From the Mailbox
Last month, in “The Data Center Became Curtailable Load,” I quoted Neil P. Osnato, founder of Persistence Analytics Group, through Data Center Knowledge. After hearing the episode, he emailed with a distinction the original report had not fully developed: a data center can be capable of curtailing electricity without being reliable enough for grid planners to count on that flexibility.
Neil examined the public PJM and Charles River Associates forms used to match large loads with new power supply. We checked them. The load form asks for projected megawatts, connection dates, ramp period, development stage, contract terms, ratings, guarantees, and credit support. The supply form asks more directly for interconnection and construction milestones, permitting, financing, land, and equipment status.
The public load-side framework does not visibly establish a standardized documentary chain proving that projected demand will arrive, ramp, and persist. That does not mean PJM, Charles River Associates, or counterparties cannot investigate those issues elsewhere. It means credit support and durable demand are not the same proof.
Neil put it this way, on the record: “Creditworthiness establishes the ability to support an obligation. It does not, by itself, establish the durability or executability of the demand that caused the obligation.”
His follow-up leaves a better question than the one we ended with: what evidence turns a projected megawatt into one the system is entitled to rely on?
Neil, thank you for listening — and for following up.
If you supervise coding or research agents, tell me how your organization counts their work. What gets called complete, how often does a person intervene, and which failed runs disappear from the productivity number? Suggested subject line: "Agent workday." Anonymous notes and source-protection requests are welcome. Email me at [email protected]. I read every message.
Listen to Episode 61
Episode 61, “The Research Org Got a Second Workforce”, is live now.
Download the episode or subscribe to the feed.
Sources
- OpenAI — Research acceleration: The view inside OpenAI
- OpenAI Research index
- Epoch AI — FrontierSWE v2
- FrontierSWE live leaderboard
- Epoch AI — Toward an O*NET for AI R&D
- Epoch AI — Contributions to OpenAI Codex show signs of AI uplift
- METR — Early-2025 AI and experienced open-source developer productivity
- arXiv — Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
- Data Center Knowledge — Fault in Data Center Alley Triggered 3 GW Load Drop
- Data Center Knowledge — PJM Says AI Data Centers Must Bring Capacity to Earn Firm Service
- PJM — CIFP Reliability Backstop Procurement / Connect & Manage
- PJM / CRA — Bilateral Request for Proposal
- PJM / CRA — Load Bilateral RFP Response Form
- PJM / CRA — Supply Bilateral RFP Response Form
Send tips, corrections, and source notes to [email protected].