To audit your agent files is to systematically inspect, evaluate, and prune persistent instruction files like CLAUDE.md and AGENTS.md so that coding agents maintain high adherence and low context overhead.
You must routinely audit your agent files to prevent performance degradation in automated software
engineering workflows. Coding agent configuration files—such as CLAUDE.md, AGENTS.md, and custom
skill packages—inevitably decay over time as underlying foundation models improve, agent runtimes
gain native capabilities, and legacy instructions accumulate.
Agent configuration files are concise operational decision guides with an expiration date, not permanent knowledge bases.
Unpruned configuration files consume finite context window tokens, trigger "context rot" through attention dilution, and frequently degrade agent adherence rather than improving execution quality. Treating these configuration files as exhaustive knowledge bases instead of concise operational decision guides wastes attention budgets on every session turn and introduces unwanted friction into your development loops.

Why coding agent configurations inevitably bloat and rot
Coding agent configurations suffer from a distinct operational half-life. Instructions written to steer older model generations or primitive runtime environments remain static while the underlying capabilities of modern foundation models rapidly advance. This dynamic creates a cognitive failure loop: every time an agent makes an error, developers respond by appending another rule to CLAUDE.md or AGENTS.md. Over time, these files balloon into exhaustive, unstructured knowledge bases rather than short decision guides. Over-specifying guidelines causes instructions to compete for model focus, leading to dropped constraints and lower overall execution reliability.

From an architectural perspective, this degradation stems directly from transformer self-attention mechanics. Transformer models process inputs by calculating pairwise token relationships across the context window with O(n²) complexity. As Anthropic's context engineering research demonstrates, "context rot" occurs as total context length increases: the model's ability to accurately recall and adhere to specific instructions degrades non-linearly because these pairwise relationships consume finite self-attention weights. Flooding the context window with static linter rules or redundant codebase summaries dilutes attention, reducing retrieval precision and reasoning performance across the entire interaction turn.
To mitigate attention scarcity, Anthropic sets an official target size of keeping CLAUDE.md files under 200 lines. Despite this threshold, unmonitored agent configurations regularly expand into hundreds or thousands of lines, wasting token budgets across every conversation turn. Demonstrating how instruction requirements shift over time, Anthropic's context engineering research revealed that removing over 80% of Claude Code's system prompt for Claude 5 generation models resulted in no measurable accuracy loss on internal coding evaluations. While this result was executed internally for specific models within a dedicated test environment and is not a universal target or public benchmark, it is clear empirical proof that instruction value expires as underlying model reasoning improves.
Empirical findings: what context files and skills actually help (and what they don't)
Empirical research across production software repositories shows that persistent context files and personalized skills rarely improve core task correctness. A controlled ablation study by Khatri evaluating 288 runs across 17 real software tasks in 3 repositories using Claude Code and Codex demonstrated that persistent context files (AGENTS.md and CLAUDE.md) did not measurably alter task correctness (bounded to ≤10–15 percentage points difference). Failure-mode triage revealed that agents fail due to deficits in core implementation skill—such as feature design, pattern selection, and exact code wiring—rather than missing repository prose that a context file could supply.

Where context files do provide measurable value is in altering agent workflow efficiency rather than execution correctness. For instance, instructing an agent via context prose that a repository's full test suite is slow prompts the model to execute targeted test suites instead of running global checks, saving session execution time. Conversely, relying on natural-language summaries of code rather than raw source text severely degrades problem-solving ability. A study by Sam-Bodden on SWE-bench Verified showed that prose summaries answered only 4 out of 45 behavioral questions about code, whereas raw source code answered 27 out of 45. Natural-language summaries smooth over implementation details; agents require direct access to raw source code to act effectively.
Similarly, developer personalization in skills yields surprisingly weak returns. An empirical study by Huang et al., analyzing 206 real-world developer sessions across 13 engineers, evaluated personalized skills extracted from developer interaction histories against generic skills pooled from broader community practices. Personalized developer skills failed to outperform generic skills because general procedural best practices transfer far more reliably across software engineering tasks. Personalized skills only added measurable value under specific conditions: when interaction histories contained multiple concrete execution examples for recurring, highly specific domain tasks (such as scheduling primitives). Stylistic or formatting preferences added noise without improving execution outcomes.
The four most common configuration smells damaging agent performance
Bad configuration patterns are pervasive across open-source repository maintenance. A repository mining study by Santos et al. (arXiv:2606.15828) analyzed 100 popular open-source repositories configured with AGENTS.md or CLAUDE.md files, establishing a formal catalog of configuration smells that degrade agent efficiency and adherence.

The most prevalent issue identified was Lint Leakage, present in 62% of analyzed files. Developers routinely dump raw linter outputs, style guides, formatting instructions, or static analysis rules directly into prose configuration files. Copying formatting guidelines into prose wastes context window tokens on instructions that are far more efficiently and deterministically enforced via automated linters, formatters, and git pre-commit hooks. Co-occurring with lint leakage is Context Bloat, affecting 42% of files. In bloated repositories, configuration files expand far past the 200-line threshold by incorporating overly detailed code examples, long architectural descriptions, redundant package lists, and duplicate information already accessible within README.md or package.json manifests.
The remaining major failure modes involve skill management and rule scoping. Skill Leakage affected 35% of repositories, where projects accumulated outdated, experimental, overlapping, or unused local and community skill packages. These orphaned skills continuously load into the agent's skill discovery budget, consuming context space on every session turn. Finally, Conflicting Instructions frequently emerged across root configuration files, nested subdirectory files, and global user scopes. When confronted with contradictory instructions across overlapping configuration scopes, language models choose between them arbitrarily, leading to unpredictable agent behavior.
The practical audit routine: doctor prompt-audit, memory review, and removal testing
Maintaining lean agent configurations requires a routine diagnostic cadence. You must distinguish between basic command-line interface diagnostics and deep session prompt auditing. Executing claude doctor in a terminal shell merely outputs local installation diagnostics and environmental status. To conduct a true configuration audit, you must run /doctor prompt-audit inside an active Claude Code session (v2.1.283+). This diagnostic command scans active CLAUDE.md, AGENTS.md, rule sets, custom skills, and tool definitions for outdated model references, missing file paths, and contradictory instructions.

# Audit a specific skill directory or run a full configuration check
/doctor prompt-audit .claude/skills/deployIn addition to prompt auditing, you must manage your skill discovery budget and persistent memory stores. Skill names and descriptions load into a discovery budget capped at 1% of the model's context window (configurable via skillListingBudgetFraction). To identify skills that consume tokens without providing value, audit skill usage statistics and flag unused skill definitions.
Also, auto-memory requires explicit review. Auto-memory is stored at ~/.claude/projects/<project-dir-name>/memory/MEMORY.md, where <project-dir-name> is derived from the Git repository name (ensuring all worktrees of a repository share a single auto-memory folder). MEMORY.md functions as an index truncated automatically at 200 lines or 25KB; Claude Code instructs the agent to offload detailed notes into separate topic files (such as user_role.md or feedback_testing.md), which are dynamically fetched on demand via tool calls rather than loaded at session launch. Execute /memory periodically to inspect and remove stale project notes, superseded developer preferences, or obsolete context from this index.
Finally, incorporate "Removal Testing" into your maintenance workflow every few weeks. Disable local skills for test runs by setting skillOverrides: "off" in your settings, or prompt the agent to execute a complex task relying solely on the raw model and base environment. If the underlying model handles the task successfully without local skills, the custom skill is a legacy crutch that should be deleted to preserve context space.
Architectural refactoring: moving rules from prose to hooks, path-scoped rules, and tests
Managing an agent's context window is fundamentally an architecture boundary problem, not a text writing exercise. Instructions written as prose in CLAUDE.md or AGENTS.md are delivered as soft contextual steering messages, which decay probabilistically under auto-compaction and long conversation histories. Conversely, PreToolUse shell hooks operate deterministically at fixed runtime lifecycle events. If a constraint must hold unconditionally—such as blocking destructive shell commands, enforcing formatting standards, or running type checks before commits—refactor the prose out of configuration files and into hard enforcement mechanisms like hooks or automated test suites.

To eliminate monolithic context bloat, refactor static CLAUDE.md files into modular, path-scoped rules within the .claude/rules/ directory. By applying YAML frontmatter containing glob patterns, rules load into the context window only when the agent reads or edits matching files, preventing irrelevant domain guidance from polluting unrelated tasks:
---
paths:
- "src/api/**/*.ts"
---
# API Development Rules
- All API handlers must validate request schemas using Zod.
- Return errors using the standard HTTP error response utility.Complement path-scoped rules with selective imports and maintainer comments. Use the @path/to/file syntax within CLAUDE.md to import external references on demand rather than duplicating text. When authoring markdown instruction files, wrap maintainer notes or human documentation in block-level HTML comments (<!-- comment -->). Claude Code automatically strips block HTML comments prior to injecting markdown content into the model's context window, though it preserves them inside code blocks or when an agent opens the file directly using the Read tool. This allows human engineers to leave maintainer notes without spending model attention tokens during automated prompts.
Make every instruction earn its place again
Context engineering is an active, continuous curation process rather than a static setup task. As coding agents, underlying foundation models, and agent runtimes rapidly mature, instructions that were vital six months ago frequently become redundant noise that actively degrades system adherence.
Adopt a zero-based approach to context hygiene: every rule, instruction, and skill defined in CLAUDE.md, AGENTS.md, or .claude/skills/ must continuously demonstrate measurable value against modern model baselines. If an instruction cannot prove its worth, refactor it into a path-scoped rule, enforce it deterministically through automated tests and shell hooks, or delete it entirely.
References
- Audit your Agent files — Addy Osmani
- Configuration Smells in AGENTS.md Files: Common Mistakes in Configuring Coding Agents — Santos et al. (arXiv:2606.15828)
- Do Context Files Help Coding Agents? A Two-Agent Ablation Study on Real Repositories — Prakhar Khatri (arXiv:2607.27250)
- Do Personalized Skills Help Coding Agents? An Empirical Study of Developer Interaction Histories — Shuyan Huang et al. (arXiv:2608.10319)
- What Context Does a Coding Agent Actually Need to Act? — Brian Sam-Bodden (arXiv:2607.09691)
- Claude Code: How Claude remembers your project — Anthropic
- Claude Code: Extend Claude with skills — Anthropic
- Effective Context Engineering for AI Agents — Anthropic