Meta’s SWE-sweep: AI coding agents need your bug report

3D rendered ai text on dark digital background

AI coding agents look a lot less capable when you stop handing them a written description of the problem. Meta’s new SWE-sweep benchmark makes that gap impossible to ignore.

What the Benchmark Found

Meta built SWE-sweep by hiding 4,068 real bugs across 100 open-source codebases. When a human provided a bug report describing the issue, the best AI coding agent resolved roughly 75% of cases. Take away that bug report and let the agent find and fix the bug on its own, and the resolution rate collapses to under 5%.

That’s not a marginal drop. That’s a 70-plus percentage point cliff.

Why This Matters for Operators

If you’re using AI coding tools in your solo or small-team workflow, this benchmark draws a clear line between what these agents actually do well and what they don’t. They’re strong at fixing a scoped, described problem. They’re still weak at the full-cycle task: notice something is wrong, figure out what, and fix it without prompting.

The practical implication is that a developer who writes precise, detailed bug reports is not just being helpful. That person is the critical variable that determines whether AI-assisted debugging works at all.

The Operator Takeaway

Don’t expect your AI coding agent to replace the judgment that goes into identifying a bug. Use it to accelerate the fix once you know what to fix. The benchmark suggests that gap is large enough that treating these tools as autonomous debuggers would be a mistake at current capability levels.

SWE-sweep is one of the more honest stress tests of AI coding performance published so far. The 4,068-bug sample across 100 codebases gives it more surface area than most benchmarks in this space.

Stay on top of AI & Automation with BizStack Newsletter
BizStack  —  Entrepreneur’s Business Stack
Logo