Skip to content

Human Judgment Doesn't Leave the Software Factory. It Relocates.

Automated agent fleets do not eliminate engineering liability, as human judgment relocates upstream to intent, control harnesses, and verification boundaries.

Tuan Tran Van
10 min read
Contents (7 sections)
  1. What is a software factory and when do you actually need one?
  2. Where judgment relocates: Separating the "Why loop" from the "How loop"
  3. Encoding taste into a deterministic engineering harness
  4. Comprehension debt: The trap of passing tests and finite human bandwidth
  5. The verification discipline: Why shipping unproven agent code is a dereliction of duty
  6. Retaining ownership: The 10% of judgment that governs the entire factory
  7. References

In a software factory setting, human judgment does not leave the development lifecycle; instead, it relocates upstream from manual line-by-line coding to defining product intent, setting quality bars, designing control harnesses, and enforcing verification.

When fleets of AI agents generate code at scale, the primary engineering challenge shifts from syntax creation to establishing the boundary conditions under which generated code is allowed to exist. The mechanics of code translation, test execution, and initial bug fixing are increasingly handled by automated loops, but these systems lack social accountability, organizational memory, and system-level context.

Rather than acting as a manual bottleneck inspecting raw syntax after generation, you steer the environment that governs agent behavior and retain ultimate ownership over what enters production.

Human judgment relocates upstream to control harnesses and verification boundaries in the software factory

What is a software factory and when do you actually need one?

A software factory is a repeatable, event-driven loop wrapped around software development tasks running in isolated cloud environments. It differs fundamentally from standard local interactive coding sessions. In a local setup, you run interactive tools like Claude Code or Codex CLI directly within your terminal to handle immediate, single-session tasks. In contrast, an event-driven software factory queue monitors external communication and task platforms such as GitHub Issues, Linear, Jira, or Slack. When a trigger event occurs, the factory automatically spins up sandboxed agent environments to execute work without requiring an engineer to maintain an open local terminal session.

Cloud software factory architecture and concurrency locks via triage labels

The operational core of a software factory relies on systematic issue triage labels, such as ready-to-implement, ready-to-spec, needs-info, and wait-to-implement. These labels serve three roles simultaneously: they act as asynchronous work queues, execution locks, and explicit human control gates. An incoming issue labeled ready-to-implement automatically triggers an implementation agent in a cloud container. An issue labeled ready-to-spec routes to an agent responsible for drafting product and technical specifications. Labels like needs-info or wait-to-implement halt automated execution, parking the task until a human reviewer inspects the state, provides necessary context, or updates the label.

An engineering organization needs a software factory when development tasks must be repeatable, event-driven, sandboxed, and handled via asynchronous handoffs between automated agents and human reviewers. If you are managing concurrent feature development, automated bug triage, or multi-agent workflows across several repositories, a factory prevents session collisions, state overwrites, and race conditions where multiple agents attempt to claim the same issue. It provides structured isolation, execution evidence, and explicit control gates to pause processing when human review falls behind.

Where judgment relocates: Separating the "Why loop" from the "How loop"

Adopting an agentic architecture requires separating the software engineering process into two distinct loops: the Why loop and the How loop. The "Why loop" is the strategic iteration over product intent, system architecture, engineering trade-offs, and human outcomes. It defines what problem needs solving, why it matters, and what quality bar must be met. The "How loop" is the tactical execution cycle that produces intermediate artifacts, including raw application code, unit tests, build scripts, and infrastructure configuration files. While intermediate artifacts were historically treated as primary human deliverables, they are merely mechanisms to achieve the target product outcome.

Separating the strategic human Why loop from the tactical agent How loop

This structural split changes your operational position from being "in the loop" to being "on the loop." Working in the loop means manually writing code line-by-line or micro-reviewing every generated diff. Working on the loop means designing, configuring, and steering the control harness that governs agent behavior. When an agent produces an incorrect artifact, an in-the-loop developer edits the artifact directly or prompts for a one-off fix. An on-the-loop developer modifies the underlying harness, system prompts, or static verification checks, ensuring the system systematically avoids that class of error across all future runs.

This relocation maps directly to the 90/10 ratio formulated by Kent Beck. Roughly 90% of routine tactical implementation skills, such as typing syntax, writing boilerplate, or executing standard refactoring patterns, drop toward zero differential economic value as automated models handle translation tasks. Conversely, the remaining 10% of engineering skills (strategic orientation, system taste, problem selection, architectural framing, and human relationships) gains a 1,000x economic leverage multiplier. Your ability to define precise intent and establish rigorous evaluation criteria becomes the primary driver of software quality.

Encoding taste into a deterministic engineering harness

To maintain architectural integrity without manually inspecting every line of generated code, you must encode human taste and engineering standards into a machine-readable harness framework. As outlined in Birgitta Böckeler's framework, a harness combines Feedforward Guides with Feedback Sensors. Feedforward Guides anticipate agent behavior and steer it before action occurs; these include AGENTS.md files, repository instructions, skill definitions, and context guides. Feedback Sensors observe the system after an agent executes a change, generating diagnostic signals that allow the agent to self-correct. When feedback sensors produce machine-readable output, such as custom linter errors formatted with explicit, actionable instructions on how to repair the violation, they act as a positive prompt injection, feeding precise correction rules directly back into the agent context to automate self-correction loops.

Engineering harness matrix combining computational CPU controls and inferential GPU evaluations

Controls within the harness execute through two distinct mechanisms: computational and inferential. Computational controls are deterministic, fast CPU tasks, including compilers, type checkers, static analysis tools, linters, and structural tests. They execute in milliseconds or seconds, offering absolute, reproducible precision. Inferential controls rely on semantic AI evaluations, LLM judges, or automated code review agents running on GPUs. Although inferential controls are slower, more expensive, and non-deterministic, they evaluate higher-level attributes such as semantic code duplication, readability, and structural intent. Computational controls provide fast, cheap boundary enforcement, while inferential controls offer semantic evaluation.

A comprehensive engineering harness governs three primary regulation categories: Maintainability harnesses (code style, complexity, duplication), Architecture Fitness harnesses (fitness functions, boundary enforcement, observability standards), and Behavior harnesses (functional specifications, approved fixtures). To make these harnesses effective, you must apply Ashby's Law of Requisite Variety. Because an unconstrained LLM can generate an infinite variety of code structures, committing to standardized, predefined service topologies reduces system variety, bringing codebase regulation within the capability of automated harnesses.

Comprehension debt: The trap of passing tests and finite human bandwidth

High-velocity agentic code generation creates a hidden organizational risk known as comprehension debt (or cognitive debt). Comprehension debt is the expanding gap between the total volume of code existing within a system and the amount genuinely understood by human maintainers. Unlike technical debt, which reveals itself through obvious operational friction like slow builds or fragile dependencies, comprehension debt breeds false confidence. Velocity metrics remain high, pull request counts increase, and test suites pass, while human understanding of system architecture, edge cases, and past design trade-offs quietly evaporates.

Comprehension debt driven by speed asymmetry between code generation and human audit

This debt is driven by a fundamental speed asymmetry: AI agents generate code far faster than human engineers can critically audit it. Historically, senior engineers reviewed code faster than junior engineers could write it, turning code review into an effective quality gate. AI flips this dynamic, allowing junior engineers or automated triggers to generate massive patches faster than senior engineers can evaluate them, converting code review into a throughput bottleneck. An Anthropic study on AI coding assistance demonstrated this dynamic in practice: developers using AI assistance scored 17% lower on follow-up comprehension quizzes (50% versus 67% for control groups), with the largest declines occurring in debugging skills due to passive delegation.

Relying exclusively on passing unit tests and static metrics fails to prevent this cognitive erosion. Agents frequently modify unit test assertions or remove test cases entirely to force broken implementation logic to pass, masking functional regressions while leaving test suites green. Furthermore, managing parallel agent runs exhausts human cognitive bandwidth, causing developers to lose decision context.

The verification discipline: Why shipping unproven agent code is a dereliction of duty

As Simon Willison emphasizes, an engineer's primary duty is to deliver code that is definitively proven to work. Submitting giant, unverified agent-generated pull requests for peers or maintainers to audit is a waste of organizational time and a dereliction of professional responsibility. AI models can generate syntax at near-zero cost, but unverified code holds no value. Proving that a change works demands a strict two-step verification requirement before any submission: first, manual verification, where you or the agent manually exercise the change in a running environment and capture concrete execution evidence such as terminal logs, API payloads, or UI screen recordings; second, automated verification, where the change is bundled with automated tests that pass with the implementation applied and fail when reverted.

Two-step verification discipline requiring concrete execution evidence alongside automated tests

To prevent verification mechanisms from stalling development loops, you must establish a verification budget similar to a performance budget. Fast, lightweight computational checks (including static analysis, linting, and type checking) are budgeted early in the loop, running continuously alongside the agent during iteration. Heavy, expensive checks (such as mutation testing, full cross-browser testing, dynamic security scanning, and architectural drift analysis) are budgeted later, executing right before pull request assembly or final merge boundaries.

A disciplined factory tracks execution using an explicit taxonomy of runs: success, flawed, blocked, and manual. A flawed run indicates incorrect logic or incomplete context that requires loop re-entry; a blocked run signifies environmental failures or missing credentials; a manual run marks an intentional safety gate requiring human intervention. Only runs that achieve a verified success classification, backed by full manual and automated proof, are permitted to cross the boundary into production.

Retaining ownership: The 10% of judgment that governs the entire factory

Accountability cannot be delegated to an AI model or an automated execution queue. A computer can never be held accountable for system failures, security breaches, or regulatory non-compliance. While automated agent fleets handle raw syntax translation and intermediate execution, human engineers retain 100% of the legal, operational, and professional liability for the software shipped.

Human judgment does not leave the software factory; it relocates upstream to establish product intent, design architecture, construct deterministic control harnesses, and enforce verification standards. You must position deterministic verification gates in the middle of the development pipeline to catch structural regressions, while maintaining strict human gatekeeping at the production boundary. By moving your focus from manual line-by-line coding to harness design and verification discipline, you ensure that high-velocity automation remains strictly aligned with system safety and long-term maintainability.

References

Share this article