A refactor that looks cleaner can quietly make your system slower, and no test will catch it.
AI assistants read code as structure, not history. A cache, a memoized function, a precomputed index: these all look like extra layers to a model optimizing for simplicity. What the model cannot see is the incident that caused someone to add that layer in the first place. That context lives in a postmortem or in someone's memory, never in the file being edited, so an AI reorganizing the code has no way to know the cache is load-bearing.
The Details:
A cache gets removed because it reads as redundant, not because anyone verified it was safe to remove. The code gets shorter and the diff looks like a cleanup. Weeks later, latency climbs and nobody connects it to a merged refactor because nothing about the change set looked like a performance edit.
Comments describing what code does get ignored by reviewers and models alike. A comment stating the measured cost of removing a cache, with a reference to the incident that justified it, changes the calculation from "looks removable" to "removing this has a known cost." That single sentence is cheap insurance against an expensive mistake.
A short, current document listing hot paths, expensive calls, and the reason each protection exists gives an assistant something concrete to check against before flattening code. It must stay accurate or it stops being trusted, by people and by the model reading it.
User-facing latency moves too late to catch this. Track the boring counters instead: cache hit rate, external calls per request, ratio of token operations to actual work. Fold these checks into the test suite so a vanished cache fails a build instead of showing up on an invoice.
Bottom Line: An AI model cannot protect knowledge that was never written down, so write it down.
Enjoy this article?
Listen to the Claude Code Conversations radio show or join the community.