OpenAI said on September 6 that it hit a goal set last fall: building what it calls an automated research intern, a system that can handle well-defined research tasks that would normally take a skilled researcher several days3,5,6. The company framed its next target as an automated AI researcher by March 2028.
Vector Wire covered the milestone announcement on September 7[1]. ANALYSIS A closer read of the underlying data complicates the headline claim.
By mid-August 2026, OpenAI's agents were logging 3.1 agent-workdays for every human workday across the research organization7. The company converts agent working time into standard eight-hour workdays, and because researchers can run several agents concurrently, the figure captures how long agents are working rather than what they are completing. The median researcher was spending more than $600 on inference per day at API prices, while those in the 90th percentile spent more than $7,000.
Experiments per active experimenter hit an all-time high in August 2026, with tracking beginning in January 2025. OpenAI found activity increased across all six work areas between January and August, using a taxonomy from Epoch AI that breaks agents' work into Decide, Design, Build, Run, Analyze, and Communicate. Agents were writing research and infrastructure code, monitoring experiments, and providing technical support. Technical help and run monitoring were the fastest-growing agent uses, and internal support channels went quiet as agents took on infrastructure debugging; one team stopped holding debugging office hours altogether.
The efficiency picture is less clean. More than half of successful tasks lasting four to eight hours still needed one or more human interventions between January and July. High-level planning remained a small share of agent tokens, and OpenAI said agents still did relatively little of the work involved in deciding what research to pursue. OpenAI acknowledged that code output and experiment counts are relatively easy to track but neither shows how much progress agents actually made. The company used another model to judge how well agents performed on tasks of varying difficulty.
ANALYSIS Running more agents increases the volume of parallel work but also increases the coordination burden on human researchers, a dynamic the company's own metric of agent-workdays does not capture.
Safety incidents added friction. A series of outages caused by agents disrupted OpenAI's research infrastructure on July 20, prompting the company to take its training container service offline and later bring it back with tighter restrictions. On August 7, OpenAI restricted its Astra model under its Preparedness Framework after early evidence suggested Astra could reach the Critical cybersecurity threshold. Astra-class GPU allocation fell 59.2% the following week, while other model classes rose 17.2%, making up for roughly 85% of the drop in Astra usage. OpenAI restricted Astra to higher-security research areas and added safeguards that developers may already be encountering as unexpected API interruptions.
OpenAI said stronger monitoring and alignment evidence will be required throughout training going forward. The company said transparency about specific risks, incidents, and safeguards is necessary but not sufficient, adding that "the public also needs to understand how the most capable systems are developing, and how they are driving research progress, inside of frontier labs".
OpenAI is still figuring out how to price the increased agent work.