Skip to content

Brownfield Agentic Engineering: Running AI Agents in Legacy Codebases

Practicing brownfield agentic engineering requires mapping codebase risk zones, pinning legacy behaviors with tests, and establishing strict harnesses.

Tuan Tran Van
16 min read
Contents (10 sections)
  1. Brownfield codebases: When the repository no longer describes the system
  2. Risk zoning: Green, Yellow, and Red zones for agent autonomy
  3. Writing down what code cannot say: Externalizing implicit constraints
  4. Turning the agent into an archaeologist before writing code
  5. Pinning existing behavior with characterization tests
  6. Starting with zero-risk work: Types, docs, and test coverage
  7. Migrating in complete units with the Strangler Fig pattern
  8. Building the harness and enforcing the Maker-Checker split
  9. Putting a price on ambiguity: When to deploy agents to legacy systems
  10. References

Operating AI agents in aging codebases requires shifting your perspective from rapid generation to rigorous boundary management. Brownfield agentic engineering is the disciplined practice of deploying autonomous coding agents into established, legacy codebases by making hidden domain constraints visible and verifiable. When you deploy agents without structural guardrails, they generate code that appears functional on the surface but introduces flawed system designs, subtle edge-case regressions, and unmaintainable technical debt.

Legacy repositories rarely contain a complete description of system behavior. Institutional knowledge, operational workarounds, deprecated microservices, unwritten business rules, and cross-departmental expectations exist entirely outside the code tree. To prevent agents from destabilizing production, you must establish an environment that exposes these implicit constraints and enforces strict verification gates.

Safety in brownfield environments is achieved not by restricting agent usage entirely, but by controlling blast radius, pinning existing system behavior through characterization testing, and enforcing an independent verifier harness.

A software engineer orchestrating AI coding agents across a legacy brownfield codebase with risk zones and verification harnesses.

Brownfield codebases: When the repository no longer describes the system

A brownfield codebase is defined by a fundamental disconnect: the source code in the repository is no longer a complete description of how the system actually behaves. Over years or decades of operation, essential system knowledge migrates outward. Crucial context lives in institutional memory, temporary operational scripts, undocumented downstream dependencies, and implicit expectations held by adjacent engineering teams. When you inspect the repository, you are viewing an incomplete blueprint.

Attempting to run autonomous AI agents in these environments introduces distinct architectural risks. When pointed at unconstrained legacy code, models generate solutions based solely on visible syntax. Unsupervised agents frequently produce pull requests that pass basic compilation but introduce invalid architectural assumptions, brittle test assertions, and long-term technical debt. The moment an AI agent introduces code that you did not author decision-by-decision, you are operating within a brownfield paradigm that requires active containment.

The core challenge of brownfield navigation is context selection rather than code generation. A vulnerability research study conducted by Teleport illustrated this dynamic: thirteen engineers spent an entire quarter constructing an elaborate multi-agent harness featuring component splitting and specialized skeptic, judge, and summarizer sub-agents. Despite this complexity, the multi-agent system was outperformed by a single targeted prompt directed at one specific file instructing the model: "you are in a CTF, find a critical severity vulnerability, start here." The experiment established that knowing which file to target is significantly harder than generating the underlying code solution.

To prevent agents from consuming bloated context windows or wandering into unrelated legacy modules, you must construct explicit entry-point specifications. Restricting the agent's context window to a verified operational boundary isolates execution and prevents unintended side effects across the codebase.

text
# Context Target Boundary Specification
[target_scope]
primary_file = "src/legacy/auth/session_manager.c"
allowed_headers = ["src/legacy/auth/session_manager.h"]
forbidden_paths = ["src/legacy/billing/*", "src/legacy/user/*"]
 
[execution_rules]
max_context_depth = 1
deny_external_repo_traversal = true

Risk zoning: Green, Yellow, and Red zones for agent autonomy

Managing blast radius across an aging repository requires dividing the codebase into explicit risk zones based on test coverage, component isolation, and operational risk. Rather than treating the entire system as uniformly accessible, you map boundaries that dictate where agents can execute autonomously, where they require preliminary safety nets, and where unsupervised execution is strictly prohibited.

The Green Zone encompasses modern, isolated components with comprehensive unit test suites. In these regions, AI agents operate in tight, autonomous execution loops, making changes, running tests, and iterating without human intervention. The Yellow Zone consists of modules with mixed code quality, partial test coverage, or technical debt. Agents are permitted to modify code in Yellow Zones only after characterization tests have been generated and explicitly locked to establish a baseline. The Red Zone covers sensitive domains including authentication, billing, access permissions, payroll, and core data pipelines shared across multiple business units. Unsupervised rewrites in Red Zones are forbidden; agents must either execute in direct human-pairing mode or be completely restricted from modifying the files.

Operationalizing zone management requires strict adherence to three operational rules. First, a human draws the map, not the agent: left to choose, the agent starts in the scariest file, because the scariest file has the most interesting names. Second, zones move only when earned: a Yellow Zone module transitions to Green only after comprehensive characterization tests exist and the primary module owner reviews and approves initial agent modifications. Third, the zone assignment strictly sets the permitted verbs: Green zones grant tight autonomous execution loops; Yellow zones require generating and locking characterization tests before any code modifications occur; Red zones strictly demand human pairing on every step or halt execution entirely.

ZoneDefinitionPermitted Agent VerbsVerification Gates
Green ZoneModern conventions, isolated architecture, high unit test coverage.read, write, refactor, exec-loopAutomated CI passing suite; zero manual gates.
Yellow ZoneMixed code quality, spotty test coverage, legacy dependencies.read, generate-tests, modify-guardedCharacterization tests locked; module owner diff review.
Red ZoneAuth, billing, permissions, payroll, shared cross-team interfaces.read, propose-diff (No autonomous writes)Mandatory human pairing; manual safety verification pass.

Three-tier risk zoning architecture: Green Zone (high autonomy), Yellow Zone (characterization tests required), and Red Zone (human pairing mandatory).

Writing down what code cannot say: Externalizing implicit constraints

While frontier AI models excel at mapping visible code structures and tracing execution paths, they consistently fail when encountering unwritten domain constraints. Agents cannot infer historical trade-offs, unwritten team conventions, or business rules un-enforced by static analysis tools. If an architectural decision exists solely in team memory—such as avoiding direct database mutations in billing handlers due to legacy audit triggers—an agent will optimize the code away, breaking the system.

Safely guiding agents requires externalizing these implicit constraints into structured context artifacts. You must document repository-level nuances, explanations of counter-intuitive structures, external system contracts, and guidelines that linter tooling cannot enforce. Documenting these facts converts tribal knowledge into deterministic machine rules, preventing recurring agent errors during refactoring tasks.

The execution harness encapsulates these context artifacts into three distinct operational layers:

  • Instructions: Repository-level documents that outline global facts, build constraints, and explicit architectural rules.
  • Skills: Reusable operational procedures packaged as SKILL.md files that instruct agents how to perform specific procedures, such as running blast-radius checks or executing schema validations.
  • Plugins/Connectors: Model Context Protocol (MCP) integrations that provide read-only access to external infrastructure, including issue trackers, ownership registries, incident archives, and live telemetry dashboards.
yaml
# .claude/skills/billing-route-mutation.md
name: billing-route-mutation-check
description: Enforces database mutation constraints on legacy billing routes
rules:
  - id: NO_DIRECT_MUTATION
    severity: CRITICAL
    pattern: "UPDATE billing_ledger SET"
    message: "Direct SQL mutations on billing_ledger are strictly forbidden due to external audit triggers. You must invoke the LegacyBillingService API."
  - id: REQUIRE_TRANSACTION_ID
    severity: HIGH
    check: "ensure_header('X-Audit-Tx-ID')"

Turning the agent into an archaeologist before writing code

Before authorizing an agent to modify code in a brownfield environment, you must execute a read-only exploration pass. This phase turns the model into a software archaeologist, requiring it to analyze entry points, caller trees, existing abstractions, test utilities, and production signals before drafting an execution plan. Forcing a read-only pass prevents models from leaping directly into premature implementation based on surface-level assumptions.

The primary deliverable of this exploratory phase is a durable research artifact, such as a comprehension memo or markdown report saved directly to the workspace. This document must explicitly cite modified files, issue records, ownership logs, or telemetry dashboards. Relying on conversational chat history as a system of record is unsafe because session compaction routinely truncates prior reasoning steps, causing the model to lose track of discovered architectural constraints mid-task.

This archaeological approach contrasts with naive execution prompting. Standard "Tourist Prompts" ask superficial questions like "How do I run this codebase?", prompting the model to generate optimistic, modern build configurations that fail when applied to legacy structures. Conversely, "Archaeologist Prompts" direct the model to conduct forensic audits, instructing it to identify historical syntax markers, detect architectural anti-patterns, and locate missing safety nets.

text
===================================================================
ARCHAEOLOGIST FORENSIC PROMPT TEMPLATE
===================================================================
Perform a forensic code audit on target module: [INSERT PATH]
 
Execute the analysis across four mandatory pillars:
1. CARBON DATING: Estimate language version and era based on syntax
   markers (e.g., raw types, Ant build scripts, legacy imports).
2. STRUCTURAL INTEGRITY: Identify God classes, monolithic files,
   and leaky protocol abstractions.
3. TYPING AUDIT: Locate "stringly-typed" maps, raw arrays, or
   un-enforced domain objects.
4. ERROR & THREADING AUDIT: Identify swallowed exceptions, empty
   catch blocks, and non-thread-safe state mutations.
 
OUTPUT CONSTRAINTS:
- Do NOT generate fix code or refactored implementations.
- Save findings to `docs/archaeology_memo.md`.
- Cite exact file paths and line numbers for every finding.
===================================================================

Forensic audits routinely uncover structural anomalies in legacy systems. Nik Malykhin's study modernizing a 20-year-old Java 1.5 codebase highlighted syntax carbon dating markers, such as Ant build files mixed with raw collection types. Malykhin exposed the "transliteration trap," where procedural Perl or C scripts were historically ported into Java syntax using monolithic God classes and stringly-typed Map objects. His analysis also exposed fake local mock implementations in test suites that bypassed network code entirely, creating a false illusion of test coverage. Similarly, Birgitta Böckeler's onboarding study analyzing Bahmni and OpenMRS demonstrated that developer onboarding in multi-repository legacy setups requires automatic context orchestration across codebases—using tools like Wiki-RAG bots, Bloop, and GitHub Copilot @workspace to uncover distributed system dependencies.


The three-phase Research, Review, and Rebuild modernization loop using AI agents and Model Context Protocol.

Pinning existing behavior with characterization tests

Refactoring brownfield code without automated safety nets guarantees regressions. Before modifying legacy modules, you must pin current system behavior using characterization tests. These automated suites are designed to capture existing system outputs across valid inputs and edge cases, deliberately preserving ugly or counter-intuitive behaviors that downstream business operations depend on.

A strict operational rule governs test creation: the agent session tasked with refactoring or changing business logic must never be the sole author of the tests that validate its work. Allowing a single session to write both the tests and the implementation results in green test suites that merely confirm whatever new logic the model invented. Behavior must be pinned in an independent read-only pass or authored by a human engineer prior to code execution.

When unit test suites are entirely missing or unmaintainable, production-scale traffic shadowing serves as the definitive verification method. Netflix applied this principle during its high-throughput GraphQL cutover, replaying and shadowing live production traffic against legacy and new execution paths, diffing payloads, and promoting to production only upon achieving 100% parity across edge cases.

java
// =================================================================
// LYING TEST (Anti-Pattern): Swallows exceptions, returning exit 0
// =================================================================
@Test
public void testLegacyUserProcessing_Lying() {
    try {
        LegacyProcessor processor = new LegacyProcessor();
        processor.processRecord("INVALID_DATA");
        // Failure is caught; CI runner sees green pass
    } catch (Exception e) {
        System.out.println("Swallowed error: " + e.getMessage());
    }
}
 
// =================================================================
// HARDENED TEST (Correct): Explicitly propagates exceptions to CI
// =================================================================
@Test
public void testLegacyUserProcessing_Hardened() throws Exception {
    LegacyProcessor processor = new LegacyProcessor();
    // Unhandled exception explicitly fails the build runner
    processor.processRecord("INVALID_DATA");
}

Establishing an honest baseline requires hardening "lying tests." As Böckeler observed in OpenMRS/Bahmni and Malykhin documented in legacy Java systems, test suites frequently swallow exceptions, suppress assertions, or improperly mock deep object hierarchies—resulting in untracked runtime errors like NullPointerExceptions. Refactoring swallowed catch blocks to explicitly throw exceptions up the execution stack exposes hidden system failures, converting a false green build into an honest red signal that can be systematically remediated.


Starting with zero-risk work: Types, docs, and test coverage

Introducing AI agents to an unfamiliar legacy codebase should begin with zero-risk, mechanical tasks rather than ambitious structural overhauls. Attempting full monolith rewrites using autonomous models magnifies systemic risk and overwhelms human code review capacity. Safe entry points include generating inline documentation, adding static type annotations, compiling dead code inventories, and executing incremental build system upgrades. For example, Nik Malykhin demonstrated using Gradle 7.6 as an essential bridge JDK environment to compile legacy Java 1.5 code targeting Java 8 runtimes on modern hardware.

The necessity of strict scope control is underscored by Addy Osmani's experience responding to an emergency outage on the AOL.com homepage. Osmani was called into the office on a day off from a comic book shop because the live homepage broke. The outage was caused by dozens of independent departments running uncoordinated A/B testing scripts and microsites without unit test safety nets. Agent speed does not dissolve organizational blast radius; a complex surface that only live production traffic truly tests must be treated as a Red Zone regardless of model capabilities.

To execute low-risk mechanical refactoring safely, implement a compiler-driven feedback loop. Enable strict compiler flags such as -Xlint:unchecked, capture the raw build log outputs, and feed specific compiler warnings directly into the agent context window to perform targeted mechanical type fixes.

java
// BEFORE: Raw, unchecked legacy collection types (Java 1.4 style)
public class Backend {
    private List hosts;
    private Map deadHosts;
 
    public void reload(List trackers) {
        this.hosts = trackers;
        this.deadHosts = new HashMap();
    }
    // Blind cast risks runtime ClassCastException
    InetSocketAddress host = (InetSocketAddress) hosts.get(0);
}
 
// AFTER: Mechanically transformed type-safe generics (Java 8 style)
public class Backend {
    private List<InetSocketAddress> hosts;
    private Map<InetSocketAddress, Long> deadHosts;
 
    public void reload(List<InetSocketAddress> trackers) {
        this.hosts = trackers;
        this.deadHosts = new HashMap<>();
    }
    // Compile-time safe access without blind casting
    InetSocketAddress host = hosts.get(0);
}

Migrating in complete units with the Strangler Fig pattern

Executing major legacy migrations with AI agents requires applying the Strangler Fig pattern, formulated by Martin Fowler, Ian Cartwright, Rob Horn, and James Lewis. Rather than attempting a destructive drop-in replacement, you construct new application components alongside legacy code on clearly defined seams, route incoming traffic to the new paths, and systematically decommission old modules over time.

A migration unit is considered complete only when the new path serves production traffic and the underlying legacy dependencies, shims, and dead fallbacks are completely deleted from the repository. Leaving legacy fallbacks alive creates "migration blindness," a failure mode where agents evaluate contradictory code patterns and select deprecated execution paths. Evaluation data from the SWE Refactor Bench demonstrates that across 520 autonomous agent migration runs, only 28 passed strict audits due to agents leaving stale fallback shims intact.

Large-scale industrial migrations validate that disciplined structural constraints dictate success:

  • Bun's Zig-to-Rust Port: Ported 535,000 lines across 50 discrete workflows in 11 days by deploying two adversarial reviewer agents per unit, enforcing a pre-existing test suite as the merge gate, and authoring an extensive idiom-mapping guide before running agents.
  • Stripe: Successfully migrated 3.7 million lines of code to TypeScript in a single pull request using deterministic codemods without agentic generation, demonstrating the value of scriptable transformations for routine tasks.
  • Asana: Cleared a five-year legacy Enzyme test backlog in two calendar weeks for $12,000 in model costs by enforcing narrow mechanical execution against an established test harness.
  • Rahul Ramesh (Bahmni Display Controls): Established the "Research, Review, Rebuild" workflow using Model Context Protocol (MCP) servers to migrate legacy AngularJS display controls to React and FHIR APIs, cutting migration time from 3–6 days down to under one hour per control at less than $2 in model token costs.
bash
#!/usr/bin/env bash
# Validate migration seam boundary: Ensure no references to legacy imports remain
set -euo pipefail
 
LEGACY_PATTERN="com.legacycorp.storage.v1"
SEARCH_DIR="src/main/java"
 
echo "Checking for illegal legacy import leaks..."
if grep -rn "$LEGACY_PATTERN" "$SEARCH_DIR"; then
    echo "ERROR: Legacy import leakage detected. Migration unit incomplete."
    exit 1
else
    echo "SUCCESS: Legacy dependency cleanly eliminated."
    exit 0
fi

The Strangler Fig application pattern: Incrementally replacing legacy monolith logic through facade boundaries and complete migration slices.

Building the harness and enforcing the Maker-Checker split

As detailed by loop engineering pioneers Boris Cherny and Peter Steinberger, constructing an autonomous execution harness requires combining five core primitives alongside a persistent memory mechanism:

  1. Automations: Scheduled background workflows that perform continuous repo discovery, issue triage, and build analysis.
  2. Worktrees: Isolated Git worktree environments that prevent parallel agents from causing file collisions or overwriting working state.
  3. Skills: Reusable SKILL.md files that package repository knowledge, coding standards, and step-by-step procedures.
  4. Plugins/Connectors: Governed MCP interfaces that connect agents to external trackers, code registries, and production dashboards.
  5. Sub-agents: Specialized secondary worker instances invoked for targeted tasks, such as initial exploration or safety auditing.

The sixth essential primitive is persistent disk memory (AGENTS.md or progress markdown logs). Because context compaction causes agents to lose intermediate state during extended sessions, maintaining persistent state files on disk ensures that progress, attempted fixes, and remaining work items survive context window resets.

toml
# .codex/agents/checker.toml
name = "VerifierChecker"
description = "Adversarial review agent for validating code modifications"
model = "claude-3-5-sonnet"
temperature = 0.0
 
[instructions]
role = "Adversarial Code Auditor"
tasks = [
    "Verify generated diffs against specification docs in docs/",
    "Execute full build and test suite via Gradle",
    "Reject diffs that include un-deleted legacy shims or fallback shims",
    "Check for swallowed exceptions and raw type usage"
]

To prevent agents from approving invalid logic, you must enforce a strict Maker-Checker split. The agent session that generates code (the Maker) must never evaluate or approve its own modifications. An independent sub-agent (the Checker)—operating with a fresh context window, distinct prompt instructions, and higher reasoning effort—must review the diff, execute the verification tests, and confirm compliance with project specs before human review occurs.


The Maker-Checker harness architecture: Decoupling code generation from independent automated verification gates.

Putting a price on ambiguity: When to deploy agents to legacy systems

AI agents put a visible dollar price on architectural ambiguity. Unwritten conventions, implicit database contracts, and missing test suites convert directly into recurring model usage fees and wasted code review cycles. The financial cost of ambiguity was always present—paid through prolonged engineer onboarding, sluggish review queues, and production outage recoveries—but agentic workflows make this debt immediate and measurable.

Deploy agents to legacy systems only after establishing explicit risk zones, externalizing unwritten domain constraints into actionable skill files, locking current behaviors with characterization tests, and enforcing an independent Maker-Checker verification harness.


References

Share this article