Harvard Study Shows AI Coding Agents Do Not Boost Output
A study by Harvard researchers reveals that AI coding agents produce significantly more code but create downstream bottlenecks, leaving overall software feature completion rates virtually unchanged.

Harvard University researchers Fiona Chen and James Stratton analyzed a dataset from Jellyfish tracking 300 million work events across more than 700,000 employees at over 700 software development firms from 2021 through March 2026. Their study shows that deploying AI coding tools creates a severe code review bottleneck, preventing companies from shipping finished software features faster despite generating substantially more code.
Deploying autonomous AI coding agents triggered a clear surge in output metrics: total lines of code rose by 30 percent, total commits increased by 20 percent, and pull requests climbed by 23 percent. Yet, these gains did not lead to a statistically significant improvement in resolving high-level Jira Issues or Epics, and researchers observed no compositional shift in the scale or complexity of those tracked items.
Instead, the efficiency of AI generation was eaten up downstream. The average length of the review process—from pull request submission to final merge—expanded by 49 percent after AI agents were introduced. The proportion of pull requests requiring changes nearly doubled, comment volume per pull request grew by 35 percent, and the share of personnel conducting code reviews increased by 14 percent. Based on Jellyfish and LinkedIn data, researchers found no significant overall employment changes caused by AI adoption.
By March 2026, 95 percent of firms had adopted AI coding agents and 80 percent used AI-based review tools. However, AI agents performed only 23.3 percent of total code review comments and submitted just 10.8 percent of all pull requests, leaving human developers burdened with filtering through machine-generated code.
This is our own summary of reporting by Ars Technica AI



