Stop AI agents from breaking your Java tests: a 6-step fix

a computer screen with a bunch of code on it

The pull request passes ./gradlew check. The diff is 400 lines. The description reads “added tests as requested.” You scroll through and find a dozen unit tests for a single service class. What you actually asked for, in a two-line Jira comment, was end-to-end coverage of the entire checkout flow.

Nobody lied. The agent filled in the gap with its most likely reading and got to work. This is the most common failure mode Java teams hit when they start using AI coding agents, and it is also the most fixable. The protocol below comes from a September 28 webinar by Sergey Pospelov of Explyt, a company that builds an AI agent for JetBrains IDEs. The author of the original write-up, Viktoria Evdokimova, notes the approach is agent-agnostic and works with Claude Code, Codex, Cursor, or an agent inside IntelliJ IDEA.

Three ways an agent burns your afternoon

One number explains the root cause of most failures: a fresh chat already spends roughly 12% of the context window on the system prompt, tool list, and any AGENTS.md file in the repository. Somewhere between 50% and 85% usage, answers get worse. At 99%, the model stops answering entirely.

  • Underspecified request. “Add tests” becomes unit tests for one class when you wanted end-to-end coverage. The fix is on your side: name the class, name the layer, reference the issue, and keep project conventions in AGENTS.md so you don’t repeat them in every chat.
  • Overfilled chat. One session, three hours, three unrelated tickets. By the fourth ticket the agent is applying a pattern it learned on the first to code from the third. One chat per task. Split large tasks across multiple chats. Compact a chat only when you’re sure the parts you need will survive it.
  • Taking the summary at face value. The agent writes “all tests pass, implementation complete” and a reviewer nods. A week later the bug is in production. Agents grade their own work generously. Verification has to be a program (tests, linter, a build) or a second agent with a fresh context, ideally from a different vendor.

As Pospelov’s slide put it: the code looked right and the agent sounded confident. Both are statements about presentation. Presentation tells you nothing about behavior.

lines of HTML codes

Step 1: Write the spec before you write the code

Solution quality, in Pospelov’s framing, is the product of how well the agent understands the project and how precisely it understands the task. Each half gets its own document.

Project document: AGENTS.md in the repository root. Build commands, module layout, test conventions, anything true for every task. Agents read it on every run, so keep it short and put it in version control.

Task document: the spec. You don’t write it from scratch. The agent, in planning mode, reads the code, asks you questions, and drafts the file. You edit and approve. If you’d rather be interviewed than face a blank page, ask the agent to ask you questions and build the spec from your answers.

Here is the spec structure from the webinar’s extended guide, rephrased for a Spring Boot rate-limiting feature:

# Task 01: rate-limit GET /api/orders

## Context
- order-service, Spring Boot 3, Java 21 (conventions in AGENTS.md)
- Origin: issue #482
- Classes involved: OrderController, ApiKeyFilter

## Goal
Cap each API key at 100 requests per minute on GET /api/orders.

## Requirements
1. Above the cap: 429 Too Many Requests plus a Retry-After header.
2. The cap is counted per API key.
3. Property orders.rate-limit.per-minute, default 100.
4. A request with no API key keeps returning 401. Do not touch that path.

## Out of scope
- Every other endpoint
- Limits shared across nodes

## Design
- RateLimitFilter registered directly after ApiKeyFilter
- In-memory token bucket per key, Bucket4j
- RateLimitProperties bound to orders.rate-limit.*

## Acceptance criteria (turn into tests before implementing)
- [ ] calls inside the cap within one minute all return 200
- [ ] the first call above the cap returns 429 with Retry-After
- [ ] two keys do not share a bucket
- [ ] overriding the property moves the cap
- [ ] ./gradlew check is green

## Plan
1. Failing tests: RateLimitFilterTest with MockMvc
2. Bucket4j dependency, RateLimitProperties
3. RateLimitFilter
4. Green, then review, three rounds at most

The “Out of scope” section stops the agent from building a distributed limiter just because it can. The acceptance criteria become the test checklist in the next step.

Step 2: Agent writes tests, you approve them, then lock the directory

The fixed order Pospelov insisted on: agent writes tests, you approve them, agent writes code, agent runs tests, repeat until green. Tests exist before the implementation does. Here is what the test class should look like for the spec above, compiled against Spring Boot 3.3:

@WebMvcTest(OrderController.class)
@Import({ApiKeyFilter.class, RateLimitFilter.class})
@EnableConfigurationProperties(RateLimitProperties.class)
@TestPropertySource(properties = "orders.rate-limit.per-minute=3")
class RateLimitFilterTest {

    static final int CAP = 3;

    @Autowired
    MockMvc mvc;

    @Test
    void callsInsideTheCapSucceed() throws Exception {
        for (int i = 0; i < CAP; i++) {
            mvc.perform(get("/api/orders").header("X-Api-Key", "key-a"))
               .andExpect(status().isOk());
        }
    }

    @Test
    void firstCallAboveTheCapIsRejectedWithRetryAfter() throws Exception {
        for (int i = 0; i < CAP; i++) {
            mvc.perform(get("/api/orders").header("X-Api-Key", "key-b"))
               .andExpect(status().isOk());
        }
        mvc.perform(get("/api/orders").header("X-Api-Key", "key-b"))
           .andExpect(status().isTooManyRequests())
           .andExpect(header().exists("Retry-After"));
    }

    @Test
    void twoKeysDoNotShareABucket() throws Exception {
        for (int i = 0; i < CAP; i++) {
            mvc.perform(get("/api/orders").header("X-Api-Key", "key-c"))
               .andExpect(status().isOk());
        }
        mvc.perform(get("/api/orders").header("X-Api-Key", "key-d"))
           .andExpect(status().isOk());
    }

    @Test
    void missingKeyStillReturnsUnauthorized() throws Exception {
        mvc.perform(get("/api/orders"))
           .andExpect(status().isUnauthorized());
    }
}

The @WebMvcTest slice registers only the controller and whatever you @Import, so both filters and the properties binding are named explicitly. Without @EnableConfigurationProperties the filter’s constructor has nothing to inject and the context fails to load. The class-level @TestPropertySource drops the cap to three, which keeps the loops short on any CI runner and turns “overriding the property moves the cap” into something the whole class exercises.

Check each test name against the acceptance criteria checklist in the spec. Four criteria, four tests, plus one pinning the requirement that must not change. Once you approve, lock the test directory so the agent can only read it. In Explyt that boundary is .agentignore; other agents use their own permission files for the same job.

⚠️ Common Pitfalls

Pospelov called out one pattern every Java reviewer should know. The agent has a failing test, can’t find the fix, and takes the cheapest path to a green build: the test file itself.

// before
assertThat(response.getStatus()).isEqualTo(429);

// after: the agent "fixed" it
assertThat(response.getStatus()).isIn(200, 429);

The isIn version reads like defensive coding. A tired reviewer will nod at it. Other shapes of the same move: a @Disabled("flaky, see follow-up") annotation that arrives with no follow-up, or a mock that quietly replaces the component under test. Build is green, acceptance criterion formally met, bug is exactly where it was.

Three defenses, in order of application:

  1. The locked test directory from Step 2. If the agent cannot write there, this entire class of shortcuts disappears.
  2. Split the roles: one chat writes the tests, a different chat implements. The implementer never sees the reasoning that produced the assertions.
  3. Route the test diff through a review agent before you read it, then read it anyway. A CI job that fails when src/test changes in the same PR as src/main without a reviewer label is cheap and catches weakened assertions on Friday afternoons.
a desk with a laptop and a potted plant on it

Step 3: Cap the review loop at three rounds

A feedback loop is any check the agent can run by itself and act on: tests, linter, compiler with warnings-as-errors, a coverage threshold. The review loop adds a second agent. The key ingredient is the hard stop. From the webinar’s example prompt:

Implement the task from task-01.md. When all tests pass, run a review subagent on your diff. If it reports problems, fix them and run the review again. Do at most 3 review rounds. Stop when the review is clean or after round 3, then report: what you fixed, what is still open, and why.

The reviewer has to be separate because the author agent will approve its own diff. The cap exists because without one the reviewer finds something on every pass, the author starts rewriting code that was fine, and the token bill grows with nothing to show for it. Anything still open after round three is a decision for a human.

✅ The full 6-step cycle

Your attention is marked in bold. Everything else the agent handles.

  1. You state the task: the issue, the goal, the limits.
  2. The agent drafts task-01.md with acceptance criteria.
  3. You approve the spec. Cheapest place to catch a misunderstanding.
  4. The agent writes failing tests. You approve them and lock the directory.
  5. The agent implements until green.
  6. The agent runs the review loop, then you read the final diff and merge.

Pospelov’s advice: escalate to this full protocol only when the simpler setup fails. The method costs real time, and it only pays off when the alternative is a rewrite.

Rules, skills, and MCP: a quick reference

Three configuration layers handle different jobs:

  • A rule is a standing instruction that goes into the system prompt on every turn. “JUnit 5 and AssertJ. Never edit generated sources.” Scope it to a file pattern so a rule about tests doesn’t follow you into production code. Rules cost tokens on every turn, so keep them short.
  • A skill is know-how loaded on demand: a Markdown file with a description the agent sees and a body it opens when the description matches the task. The open Agent Skills format lets agents from different vendors share skill files, so a team using two agents doesn’t maintain two copies.
  • An MCP server is the agent’s reach outside the IDE: GitHub, Jira, a browser. If a server exposes fifty tools, their descriptions alone eat context. Some agents can park such a server behind a subagent so the main chat only sees a summary.

Pospelov’s rule for choosing: if it should always apply, write a rule; if it’s a procedure for some tasks, write a skill; if it needs data or actions outside the editor, connect an MCP server.

️ The reviewer’s checklist

Print this and keep it next to your monitor:

  • Does the spec name the classes, the issue, and what is out of scope?
  • Does every acceptance criterion have a test, and does every test trace to a criterion or a requirement?
  • Is the test directory read-only for the agent during implementation?
  • Did the review loop have a round limit, and what was still open when it stopped?
  • Did any assertion get weaker, any test get disabled, any mock replace the component under test?
  • Can you explain the diff to a colleague without the agent’s summary?

The full spec template, slide deck, and prompts from the September 28 webinar are in Explyt’s extended guide. The rate-limiting example uses Bucket4j for the token bucket implementation. Parallel task work across unrelated tickets is best handled with git worktree, separate checkouts on separate branches, rather than multi-agent pipelines, which rebuild context from zero on every subagent and cost noticeably more than a single focused agent doing the same job.

Stay on top of AI & Automation with BizStack Newsletter
BizStack  —  Entrepreneur’s Business Stack
Logo