On September 6, 2026, OpenAI said its internal measurements show it has reached its automated research-intern goal: handling well-defined research tasks under human direction, including work that might take a skilled researcher days. The announcement reports internal progress and usage, not the launch of a service named “research intern” or a claim that AI independently chooses research priorities.
What do 3.1 agent-workdays actually measure?
By mid-August, OpenAI says its research organization used 3.1 agent-workdays per human workday, with both expressed as eight-hour days. This compares accumulated runtime, not research quality or output. Parallel sessions add hours without necessarily shortening the time needed to solve a particular problem.
The clearer change is task allocation. Alongside research and infrastructure coding, agents increasingly help with troubleshooting and monitoring experiments. High-level planning remains a small share of output tokens. The report therefore describes substantial delegation of execution, not the disappearance of human decisions about questions and findings.
More agent use is not the same as more discovery
Usage at the median researcher exceeded $600 a day and at the 90th percentile exceeded $7,000 a day, valued at API prices. These are inference-usage valuations, not disclosed internal bills or the daily price of a consumer Codex subscription.
Experiments per active experimenter also rose, but available compute grew too. Correlation alone cannot assign all gains to agents. More code and more runs are different measurements from reliable, reproducible discoveries.
Longer tasks still depend on human steering
Over the past six months, more than half of successful tasks estimated to take a human four to eight hours involved at least one intervention. That duration estimates human task difficulty, not agent runtime. Uncertain outcomes were excluded, so the figures do not cover every task.
An automated AI researcher by March 2028 remains a goal, not a completed result. OpenAI says it does not yet know how to safely reach full recursive self-improvement. People still choose priorities, evaluate which results to pursue and decide whether to scale, pause or deploy.
This is not a promise of one-click research
For teams and creators, the useful distinction is between a task with a clear deliverable and an open-ended assignment with no success criterion. Runtime alone cannot establish completion. The questions are what was delivered, how much correction people supplied, and whether the result can be trusted.
