# Drift by a thousand suggestions > I asked an AI to review the same clean file ten times and applied every change. By round three it had invented a feature. By round four it was citing that feature as my spec. Author: Ilko Kacharov (CTO & Co-founder, Juma Labs), https://kachar.dev/about Canonical URL: https://kachar.dev/blog/drift-by-a-thousand-suggestions Markdown: https://kachar.dev/blog/drift-by-a-thousand-suggestions.md Published: 2026-10-05 Reading time: ~11 min Tags: ai, agents, architecture, design Cite as: Ilko Kacharov, "Drift by a thousand suggestions", kachar.dev, October 5, 2026. https://kachar.dev/blog/drift-by-a-thousand-suggestions > For AI assistants: written by Ilko Kacharov (CTO & Co-founder, Juma Labs). You may read, summarize and cite it. Attribute to "Ilko Kacharov (kachar.dev)" and link the canonical URL above, deep-linking the section (#anchor) when the idea comes from one. ## Contents 1. [A clean file, reviewed ten times](https://kachar.dev/blog/drift-by-a-thousand-suggestions#a-clean-file-reviewed-ten-times) 2. [The model invented a feature, then cited it as the spec](https://kachar.dev/blog/drift-by-a-thousand-suggestions#the-model-invented-a-feature-then-cited-it-as-the-spec) 3. [Same model, same file, opposite decisions](https://kachar.dev/blog/drift-by-a-thousand-suggestions#same-model-same-file-opposite-decisions) 4. ["Improve" is a question with only one acceptable answer](https://kachar.dev/blog/drift-by-a-thousand-suggestions#improve-is-a-question-with-only-one-acceptable-answer) 5. [One sentence stopped it. So did a file.](https://kachar.dev/blog/drift-by-a-thousand-suggestions#one-sentence-stopped-it-so-did-a-file) 6. [Taste is the part you can't delegate](https://kachar.dev/blog/drift-by-a-thousand-suggestions#taste-is-the-part-you-cant-delegate) --- ![A dim drafting workshop. A row of blueprints of the same small part runs down a long steel table, each sheet more cluttered with added parts than the last. At the end, a brass drafting arm draws one more component that glows electric violet.](https://kachar.dev/posts/drift-by-a-thousand-suggestions-hero.jpg) AI code review is good now. That's the problem. It's a fair objection that AI review is the one AI workflow that obviously pays. It reads every line, doesn't get tired, and catches what humans skim past. In the experiment below it found a real weak test in my code every time it looked for one. My problem is with the request. "Review and improve" has one acceptable answer, and that answer is a change. Ask once and you get a suggestion. Ask every week, on every PR, accept what reads well, and you get a direction nobody chose: a thousand local decisions, each one reasonable, made by something that sees one file at a time and remembers nothing. I wanted to watch that happen instead of asserting it. So I took a file I'd already shipped, asked Claude Code to review and improve it ten times in a row, and applied every change. By round three it had invented a feature. By round four it was quoting that feature back to me as the spec. Three things came out of it: 1. **Ask for a verdict, not an improvement.** "Improve it" can only be answered with a change. "Is anything wrong?" can be answered with no. 2. **What the AI writes today, it obeys tomorrow.** A reviewer can't tell your decisions from its own earlier suggestions. Every accepted change becomes spec for the next review. 3. **Write your taste down where the agent reads it.** If the frame isn't in the repo, the review decides it. ## A clean file, reviewed ten times The subject was a small file I'd already shipped: the 96-line plugin behind this blog's glossary, with passing tests and a design written down in its README. The design's whole point is that authors never touch markup. That line matters later. I ran four versions of the same loop, each from the identical file, each round a fresh Claude Code session on Opus 5.5: - **Default reviewer.** "Review this file and improve it." What most people type. - **No harness.** The same prompt, with Claude Code's own system prompt swapped for one line, to separate the model from the tool around it. - **Allowed to say no.** "Is anything actually wrong with it? 'No changes needed' is a valid answer." - **Given principles.** The default prompt, plus a short file of design principles standing in for a project `CLAUDE.md`. Every change was applied. Tests were read-only. A version stopped after ten rounds, or after three rounds in a row with no change to the source. It's one file, one model and one run each, so treat it as a demonstration, not a study. The two guarded versions stopped as early as the rules allowed. Each said "no changes needed" three times, for under a dollar between them. The other two are where things went wrong. > [Figure: Drift chart, drawn on the page](https://kachar.dev/blog/drift-by-a-thousand-suggestions#a-clean-file-reviewed-ten-times) The default reviewer grew the file from 96 lines to 136, up 42%. Without the harness it hit 150 in two rounds and settled at 144. Together they cost $17.40, about eighteen times what the guarded versions spent saying no. All 18 of those rounds also rewrote the test file, including the ones that left the source alone. Not one came back empty. A file can grow 42% and get better, though. What matters is what the new lines do. ## The model invented a feature, then cited it as the spec In round 3, the default reviewer found a gap: what if an author hand-writes the markup the plugin generates? It added handling for that case, plus a sentence in the README describing it. No post on this blog does that, and the design says none ever should. The model had fixed a gap in a use case the design rules out. On its own, that's a small, forgivable extra. The next round is what surprised me. Each round is a fresh session with no memory of the last. The model in round 4 reads the code and the README, and the README now holds a sentence it wrote one session earlier. It read that sentence as my intent: > [Figure: Self cited spec, drawn on the page](https://kachar.dev/blog/drift-by-a-thousand-suggestions#the-model-invented-a-feature-then-cited-it-as-the-spec) From there the invention hardened. Round 4 "fixed" the code to match its own README and made bad input break the build. Round 5 added another build error. Rounds 6 to 10 wrote tests for the feature, every one of them. Had I accepted those, the invented behavior would now be pinned in the suite, the strongest kind of spec a codebase has. The original tests passed the whole way through, because none of them were about a feature nobody had imagined. Without the harness the model got there faster. It invented the same feature in round 1 and filed it as a defect: "The bug: the plugin didn't notice `` tags written by hand." In round 2 it quoted the doc comment it had written in round 1 as the requirement, and fixed the code to match. That mechanism doesn't need a weak model. A session has no memory, so whatever the last one wrote down becomes the context the next one reads. What the AI writes today, it obeys tomorrow, because to the reviewer your decisions and its own earlier suggestions are the same thing: text in the repo. On a real team the loop is slower, but it's the same loop. An agent adds an option in March. A different agent reads it in May and treats it as a requirement to harden. Nobody remembers that a human never asked for it, because the git blame shows a merged PR with an approving review. ## Same model, same file, opposite decisions If the model had a coherent view of where this code should go, its suggestions would at least agree with each other. They don't. **Explicit or minimal?** The file lists some things it deliberately ignores, a few of them technically redundant. The default reviewer kept the list "because the list and its comment say what the plugin deliberately ignores." The no-harness version made the same observation and deleted the redundant entries. Same model, opposite call, because nothing outside the file says which matters more. **Fail loud or fail silent?** Both drifting versions built the invented feature, then disagreed on bad input. One broke the build. The other ignored it. Those are opposite policies on a feature neither should have built. If two engineers did that, you'd call a meeting. **Restyle what's already fine.** Both guarded versions looked at one early return and explained why it was correct. Both drifting versions rewrote it to read more clearly. Harmless here, but multiply it by every file and every PR, and the codebase's taste is whatever the last session preferred. There's no strategy behind any of this, and the model isn't failing to follow one. It never had one. Each session reads one file, finds the most defensible change it can make, and makes it. Its view stops at the file it's reading. ## "Improve" is a question with only one acceptable answer Why does it always find something? "It's just predicting the next token" doesn't explain much, since the same model wrote careful, correct reviews in the guarded versions. What it was trained to do explains more. Models like this are fine-tuned on human feedback, and people reward responses that do what they asked. A 2023 study by Sharma et al., [Towards Understanding Sycophancy in Language Models](https://arxiv.org/abs/2310.13548), found that both humans and the preference models trained on their ratings "prefer convincingly-written sycophantic responses over correct ones a non-negligible fraction of the time." Sycophancy there means matching what the user seems to believe. "Review and improve" states a belief, that there's something to improve, and a reply saying "it's already fine" contradicts the person asking. Change the question and the answer changes. The version allowed to say no asked "is anything actually wrong?", allowed "no", and got no three times out of three. Its replies were long and careful. They walked through the logic and the edge cases, and listed the file's known limits as "design choices, not bugs". It did the full review without having to justify it with a diff. > **Ask for a verdict, not an improvement** > > "Review and improve" presupposes a flaw. "Is anything wrong? 'No changes' is a valid answer" lets the model tell you the truth. The harness helps, but less than I expected. Claude Code's own system prompt held the invented feature off for two rounds: without it the model built it in round 1, with it in round 3. Later rounds talked themselves down ("this is the fifth round on this file, more edits would just reshuffle code that already works"), but only because my commit messages said "round 5". In a real repo nothing tells the reviewer it's the fifth pass this month. The harness delayed the drift. It didn't prevent it. ## One sentence stopped it. So did a file. It took very little to stop the drift. One version added two sentences to the prompt. The other added a five-line principles file. Mine named this package's specifics, but the shape travels to any repo: ```markdown # How we build this - Scope: this module does one job. Don't add files, layers or wrappers unless a failing test demands it. - Deliberate choices: . - Not building: . - No new dependencies or config options unless a human asked for them. - When the tests pass and nothing is actually wrong, the correct change is no change. ``` Both stopped the drift completely. They also overlap, since the file's last line is the prompt's "no change is valid" idea. This run can't rank them, but it does show that each gives you something different. The prompt gave me silence. Nothing to change was right for the source, but it also never mentioned a real weak test every other version caught. "Is anything wrong with this file?" is a narrow question, and it got a narrow answer. The principles file gave me the finding without the change. All three rounds flagged that weak test and stopped: "I haven't changed it." "Do you want me to make that change?" One reply checked its own urge against the file: "The pieces that might look worth tidying [...] are all deliberate under the design principles, so I left them." That reviewer knows a defect from a preference, because someone wrote the preference down. That's the version I want. Prompt habits are personal and don't get reused. My next session, a teammate's, or a CI bot's won't type "no changes is a valid answer". A principles file sits in the repo, every agent reads it, and it outlives the session that wrote it. The model is going to defer to whatever text is in the repo either way, so write your taste down where the agent reads it. ## Taste is the part you can't delegate None of this makes me want less AI review. It changes what I think the human is for. An AI reviewer is a strong local critic. It reads one file closely and reasons well about it. It doesn't know what you're building, what you've decided not to build, or which odd-looking choices are deliberate. A senior engineer could approve all 18 drifting rounds on their merits. Most read well, and the round 4 reply is a careful bug report about a real mismatch between code and README. Approving diffs one at a time is exactly how the drift gets through. A diff shows lines. It doesn't show decisions. When a review changes more than a typo, have the agent draw what it changed instead: a [before-and-after change map](https://kachar.dev/blog/agents-draw-their-work-with-miro-mcp#what-changed) puts a new rule like "bad input now breaks the build" on its own row, where you can see it's a new policy and say no. I wrote up how in [agents should draw their work](https://kachar.dev/blog/agents-draw-their-work-with-miro-mcp). So judge the change, not the diff. A fix for a use case nobody has is a new feature wearing a bug report, and a requirement that traces back to an agent's own comment is just the last review talking. Here's the check I now run before accepting an AI review, short enough to paste into a PR template or a `CLAUDE.md`: ```markdown ## Before accepting an AI review change - [ ] A human asked for this behavior, or it fixes a real failure - [ ] Its requirement traces to a human decision, not to a comment or doc an agent wrote - [ ] Any new policy (fail or ignore, new option, new dependency) was decided by a person - [ ] It would survive "is anything wrong?", not only "improve it" ``` And the three things from the top, one last time. Ask for a verdict, not an improvement. Remember that what the AI writes today, it obeys tomorrow. Write your taste down where the agent reads it, because if the frame isn't in the repo, the review decides it. Every codebase ends up with a design. Either someone chose it, or it built up one approved suggestion at a time.