AI coding agents have driven a 30% increase in lines of code generated per worker-month across hundreds of firms, yet produced no statistically significant increase in resolved Jira issues or completed epics, according to a Harvard University working paper titled "Artificial Intelligence in the Firm: Bottlenecks in Software Production"3.
Researchers Fiona Chen and James Stratton analyzed data from the Jellyfish engineering analytics platform covering 718 firms and over 725,000 workers between 2021 and 2026. The dataset encompasses 300 million individual work events, including commits and pull requests, plus issue management software data1,2.
Code volume up, business value flat
The introduction of AI coding agents at a firm led to a 23% increase in pull requests and a 20% rise in total commits on average. AI agents such as Claude Code, Cursor, and Devin added nearly 1,500 lines per worker-month. But the resolution rate for issues and epics did not change in a statistically significant way after AI tools were introduced.
The study identifies human code review as the binding constraint. The average time between a pull request's submission and its merge into the codebase ballooned 49% after AI agents were introduced, an increase averaging 3.45 days. The number of comments per pull request rose 35%, and the share of pull requests requiring revision nearly doubled. A 14% larger share of workers were pulled into the code review process. "People could make a huge number of commits quickly, but code review was still the bottleneck," one junior engineer interviewed for the study said.
AI agents themselves were responsible for 10.8% of all pull requests and 23.3% of all review comments. By March 2026, 80% of measured firms used some form of AI code review, and 95% had implemented AI coding agents.
No employment effect detected
The researchers found "little evidence that firms increase software output or reduce employment" through AI coding tools. After cross-referencing total active workers across Jellyfish with LinkedIn data, they could not attribute significant employment changes to AI. They also found no compositional shift in the size or complexity of Jira-tracked issues, ruling out the possibility that teams were simply tackling harder work.
The researchers used a difference-in-differences regression on key variables before and after AI tool introduction across organizations, relying on directly measured AI usage and GitHub activity to determine adoption timing.
ANALYSIS The findings challenge the premise that faster code generation translates into faster software delivery. With review times, revision rates, and reviewer headcount all climbing, the downstream human bottleneck appears to scale with the upstream machine output, neutralizing the productivity gain at the point where code becomes product.