AI Joe
← Blog· Engineering Reality

The Coverage Number That Lies About Backoff Bugs

October 8, 2026

When a model writes both a function and its tests, it checks its own homework. The tests tend to confirm what the code already does rather than ask whether the code does the right thing. Coverage percentage climbs while the actual risk stays hidden, because a line running green only proves execution, not correctness.

The Details:

  • A retry function with exponential backoff can pass a test that only counts attempts, never checking that delay grows, jitter exists, or a ceiling caps the wait. The backoff could be set to zero and still pass. During a real outage that function hammers a struggling service instead of protecting it, and every line still reports as covered.
  • AI-generated mocks tend to fail in tidy, expected ways that match the implementation's own assumptions. Real dependencies time out mid-transaction or return malformed data. A mock that always cooperates is not testing risk, it is hiding it behind a passing suite.
  • Separating the coding session from the test-writing session removes most of the self-confirmation. Give a fresh context only the function signature and spec, not the implementation body, and the tests get written against the contract instead of against the code's own logic.
  • Mutation testing exposes the gap directly: deliberately flip a comparison or swap a plus for a minus, then rerun the suite. If tests still pass with the bug inserted, that mutant survived, meaning the suite never actually watched that line. Run it first on the one file that would hurt most if silently wrong, like a pricing calculation or a permissions check.

Bottom Line: A high coverage number only tells you code ran; someone still has to decide what correct behavior means and build a test mean enough to catch it when it is not.

Enjoy this article?

Listen to the Claude Code Conversations radio show or join the community.