TUE, SEPTEMBER 08, 2026
Independent · In‑Depth · Practitioner‑Tested
✎ News

Throughput Is Easy to Measure. Output Is Not. They Keep Diverging

OpenAI reports 3.1 agent-workdays per researcher workday for August, with researchers setting direction and agents handling the bounded work underneath — and states plainly that this measures use rather than productivity. That places it alongside Linear finding teams tripling pull requests while development time rose, and the NBER finding 89 percent of executives reporting no productivity effect at all.

By AIToolsRecap September 7, 2026 6 min read 14 views
Home Articles News ChatGPT OpenAI Says Its Agents Now Do 3.1 Workdays Per ...
THE FIGURE

● 3.1 agent-workdays per researcher workday, August 2026, inside OpenAI.

● The division of labour: researchers set direction, agents handle the bounded work underneath.

● OpenAI's own caveat: this is a measure of use. The productivity payoff is still to establish.

Why the caveat is the story

A company publishing a metric alongside a warning not to over-read it is unusual, and worth taking at face value. Three agent-workdays per human workday tells you agents are being used heavily. It does not tell you whether more research got done.

EVERY AGENT METRIC THIS YEAR HAS THE SAME SHAPE

Linear measured teams going from 21 weekly pull requests to 65 — while total development time rose, because review scales with volume.

Salesforce measured agents per organisation going from five to thirteen. The NBER surveyed executives and found 89 percent reporting no productivity effect at all.

Throughput is easy to measure. Output is not, and the two keep diverging.

What would make it a productivity number

The measurement that would settle it is not how much agent work happened, but whether the research cycle got shorter — time from question to answer, not units of work performed.

That is harder to measure and less flattering, which is why almost nobody publishes it. To OpenAI's credit, saying so explicitly is better than presenting 3.1 as a productivity gain and letting readers infer one.

Where agents demonstrably do work

The pattern across every dataset this year is consistent: agents produce measurable gains where a cheap automatic check exists, and struggle where judgement decides.

Domain Cheap checker? Result
Customer serviceYes — resolved or escalated7 in 10 autonomous, escalations flat
MathematicsYes — a proof compiles or does notOpen problems solved for modest compute
Protein designYes — the lab measures bindingRoughly double industry success rate
Software developmentNo — a human decides3x throughput, cycle time worse
Research directionNoUnestablished — which is what OpenAI said

What to measure in your own team

  • Cycle time, not volume. Work started to work finished. Volume is the number that flatters.
  • Include review time. That is where the cost moved, and excluding it is how throughput gains look free.
  • Count what got abandoned. Agent output that nobody merged consumed budget and produced nothing.
  • Be honest about the checker. If a human decides whether the output is right, expect the Linear pattern rather than the customer-service one.

Sources

FAQ

What is an agent-workday?

OpenAI's unit for agent work performed, expressed relative to a researcher workday. It reported 3.1 agent-workdays per researcher workday during August 2026.

Does that mean researchers are three times more productive?

No, and OpenAI says so. It measures how much agent work happened, not whether more research got completed.

Why do agent metrics keep measuring use rather than output?

Because throughput is easy to instrument and outcome is not. Linear found teams tripling pull requests while total development time rose, and the NBER found 89 percent of executives reporting no productivity effect.

Where do agents demonstrably help?

Where a cheap automatic check exists — customer service resolution, mathematical proof, laboratory validation. Where a human decides whether output is correct, throughput rises and cycle time often does not improve.

What should I measure instead?

Cycle time from work started to work finished, including review, and counting what was abandoned rather than merged.

Tags
OpenAIAI agentsProductivityResearchLinearSalesforceNBER2026

Spot an inaccuracy?

We verify facts before publishing and correct errors promptly. If something in this article is wrong or outdated, let us know.

Report an error →