What Belongs in AGENTS.md, Design Docs and Tests?

Why agent instructions accumulate and how to decide what to keep, move, or remove.

·
·
What Belongs in AGENTS.md, Design Docs and Tests?
Read3 min
  • The study tracks 247,694 instruction lifetimes across 1,867 repositories and reports average instruction-count growth of 226%.
  • In one controlled experiment, background notes reduce excess instructions by 99.3% compared with no comments, at similar constraint satisfaction.
  • The author recommends a short instruction file with essential guidance and pointers to detailed sources.
  • Design documents explain architectural decisions. Tests verify required behavior, while tools and hooks can enforce restrictions.
  • Preserve each rule’s original failure, rationale, and supporting evidence so maintainers can assess future changes.

Every time an agent makes a mistake, you face a maintenance decision. You can either add an instruction to clarify a decision, or write a check that detects the fail case. The choice determines where the rule for that requirement is enforced and what you have to maintain down the road.

An instruction file like Claude.md can tell the agent how to work with a repository. A design document can explain why a module behaves a certain way. A test can check whether a change preserves required behavior. A single requirement may need all three, with each being responsible for a different part of the decision.

The difficulty comes when the instruction file inherits every correction and nobody revisits those decisions. The paper “Why Does CLAUDE.md Keep Growing? Catastrophic Remembering in Agentic Coding” examines that maintenance problem. Its author, Kushal Chakrabarti, extends the research into a practical framework in comments to AlphaSignal.

“Treat your instruction file like a best-effort cache,” Kushal said. “Fill it on a miss, refactor regularly to a source of truth, and evict aggressively.”

That framework gives you a way to evaluate the next rule before it becomes a permanent part of every task.

Why Instruction Files Keep Their Old Rules

The paper tracks 247,694 instruction lifetimes across 1,867 repositories. It reports average instruction-count growth of 226% over tracked file lifetimes. While engineers occasionally do manual cleanups to prune the text, these rewrites only offer temporary relief. Without an eviction policy, teams quickly resume appending ad-hoc rules until the bloat returns.

This pattern recurs because rules outlive their context. When an instruction survives but the original failure is forgotten, future maintainers can’t tell whether it protects against an active bug or patches a long-resolved quirk.

Let’s consider a hypothetical Python service with this instruction:

code
Do not upgrade the export dependency beyond version 2.4.

An earlier upgrade causes the exporter to omit the final record from some batches. Someone adds the rule, but the instruction contains no link to failure or regression test.

Later, when a new dependency version becomes available, the agent still sees the prohibition. A maintainer cannot tell whether it protects against an unresolved bug or preserves a workaround that no longer applies.

The maintenance problem still remains. The rules still state a decision without the evidence needed to reconsider it.

This distinction matters when you audit instructions. When you set a word- or line-count limit on your instructions file, you only specify the target length. You do not specify which requirements still apply or what might break after an edit.

Preserve the Evidence Behind the Instruction

To test whether preserved context helps maintainers revise instructions, the paper assigns two roles to language models. The maintainer updates a set of rules based on feedback from task attempts. The executor uses those rules to complete instruction-following tasks. In the experimental setup, the maintainer must infer the requirements from feedback because it cannot see the full reference instructions.

The researchers compare maintainers with and without access to background notes, which the paper calls informative comments. These notes explain which failure prompts an instruction, why the maintainer expects it to help, and what previous attempts reveal. The executor receives only the rules. This separation tests whether the notes help the maintainer decide what to keep or remove.

Keep reading

Don't miss what's next in AI

Join 300,000+ engineers and researchers who get the signal, not the noise. Create a free account to read the rest of this story.

  • Full access to in-depth AI research breakdowns
  • Be the first to know what's trending before it hits mainstream
  • Daily curated papers, repos, and industry moves
Trending
  • No trending articles

Comments

avatar

Next Reads