More compute for a single AI coding agent produces modest gains in bug detection. A second agent reviewing the first one’s work produces much larger gains. That is the core finding from new research tied to Meta.
Meta’s internal numbers
Meta’s automated review tool, RADAR, has processed over 535,000 diffs. More than 331,000 of those were merged into the codebase. RADAR cut median review wall time by 35%. With risk calibration active, Meta reports a lower revert rate than manual review achieved, and production incidents dropped to one-fiftieth of the rate seen under manual processes.
Meta also ran its Engineering Agent on test failure repairs over a three-month trial. Of the fixes it generated, 80% went through human review. Approximately 25.5% of those reviewed fixes made it into production.
Mutation testing at scale
Meta’s Just-in-Time testing framework combines large language models, program analysis, and mutation testing. Mutation testing works by planting deliberate faults in code to check whether your test suite catches them. Across more than 22,000 generated tests, the approach produced a fourfold increase in bug detection overall. For significant failures, the improvement reached up to twentyfold.
Agents checking agents
A protocol called Adversarial Review, tested on the LiveCodeBench benchmark, hit an 87% pass rate. That beat single-agent setups and some configurations using five agents. A separate system called Wink, focused on catching coding agents that go off course, recovered from approximately 90% of misbehaviors across more than 10,000 real-world instances.
The shift for engineering teams
The Engineering Agent data points to a changing role for engineers. Roughly three-quarters of machine-written fixes did not make production, which means the job is moving from writing every fix to judging which AI-generated fixes deserve to ship.
