If you’re letting AI agents merge code into production, a new study puts a number on how often that code needs a follow-up fix. The answer is roughly 62 percent more often than equivalent human-authored PRs from the same repositories over the same time period.
What the Research Measured
Researchers tracked 6,774 merged agent pull requests across five tools: OpenAI Codex, GitHub Copilot, Devin, Cursor, and Claude Code. They compared them against 5,044 contemporaneous human PRs from the same open-source repositories (all with at least 500 GitHub stars). Every candidate fix was verified by human annotators and an LLM judge that matched human-level agreement, with a direct-fix precision of 90 percent.
The Three Key Findings
- Merged agent PRs attract verified fixes at 1.62 times the odds of merged human PRs from the same repos over the same window.
- 69.6% of those fixes come from the same agent that authored the original PR.
- 76.4% of verified fix PRs are agent-authored throughout all commits, meaning humans are rarely the ones cleaning up.
The headline summary from the researchers: agents currently largely finish their own job, but their merges still require fixing more often than human merges.
The Operator Takeaway
If you’re using any of these five agents in a real codebase, the data suggests you should not treat a successful merge as the end of the story. Budget for follow-up passes. The reassuring part is that the same agent will handle most of those fixes without human intervention. The less reassuring part is that the loop exists at a meaningfully higher rate than it does with human contributors.
