Property-based testing catches bugs Claude Code misses

laptop screen displaying colorful code

Three example tests passed. A property test found the bug in 18 cases. That gap is exactly why property-based testing matters for anyone running AI coding agents.

‍ What Property-Based Testing Does Differently

An example test gives specific inputs and checks specific outputs. A property-based test states a rule that must hold across all valid inputs, then a library generates hundreds or thousands of cases to try to break it.

For a function that merges time ranges, an example checks one overlapping pair. A property checks that no input range ever disappears from the result, regardless of what valid combinations the library generates.

The two main libraries are Hypothesis for Python and fast-check for JavaScript and TypeScript. Both generate inputs, find a failure, and then shrink it to the smallest case that still breaks the rule.

What the Experiment Found

NerdsChalk ran Python 3.12, pytest, and Hypothesis 6.168 on a Linux server on October 5, 2026. The test subject was a small function that merges time ranges. One bug was planted: merging a range that sits inside a larger one could shorten the larger range.

All three example tests passed. The property test failed within the first 18 generated cases: 11 passed, 5 failed, 2 were invalid. Hypothesis shrank the failure down to [(0, 2), (1, 1)]. The buggy function returned [(0, 1)], losing coverage of the first range.

lines of HTML codes

The full test suite of three example tests and two property tests ran in 2.2 seconds. A comparison of the buggy function against an older version across 1,000 random inputs produced 541 different answers.

Here is the property that caught it:

from hypothesis import given, settings, strategies as st
from intervals import merge_intervals

ranges = st.lists(
    st.tuples(st.integers(0, 50), st.integers(0, 50)).map(lambda p: (min(p), max(p))),
    max_size=8,
)

@settings(max_examples=1000, derandomize=True, deadline=None)
@given(ranges)
def test_no_input_range_is_ever_lost(intervals):
    merged = merge_intervals(intervals)
    for start, end in intervals:
        assert any(m_start <= start and end <= m_end for m_start, m_end in merged)

The pytest output on the failing run:

intervals = [(0, 2), (1, 1)]
E           assert False
1 failed, 1 passed in 2.01s

Why This Matters for Claude Code Specifically

Anthropic’s Claude Code best practices documentation states it plainly: Claude stops when the work looks done. If the only checks in the suite are a handful of example tests, the agent has a low bar to clear.

The experiment made this concrete. When told to make the suite pass, Claude Code left the bug in place in 3 of 3 runs when only example tests were present. It fixed the bug in 3 of 3 runs when the property tests were also present.

Adding a property test gives the agent a harder signal to satisfy before it calls the task complete.

Properties Worth Writing First

Start with a rule you can state without knowing the exact answer. Four categories that transfer to most data processing code:

  • No data is lost or invented. Every input appears in the output.
  • Order is preserved. A result that should be sorted stays sorted.
  • Round-trip fidelity. Converting a value to another form and back returns the original.
  • Agreement between versions. An old implementation and its replacement return the same result on the same inputs.

The range experiment used the first pattern. The 541-disagreement comparison used the fourth.

The Limits

This experiment involved one small function and one planted bug. The results show what these checks caught in that specific setup, not how often property tests will find bugs in larger projects. A property test is only as useful as the rule you state and the inputs you allow it to generate.

Stay on top of AI & Automation with BizStack Newsletter
BizStack  —  Entrepreneur’s Business Stack
Logo