Skip to content

Graph Engineering: The Karpathy Loop and Anthropic's Playbook

Graph engineering transitions autonomous agent swarms from volatile single-agent iteration loops into durable, graph-grounded execution systems.

Tuan Tran Van
14 min read
Contents (9 sections)
  1. Autoresearch and Karpathy's Self-Optimizing Loop
  2. From a Single Loop to AgentHub: Git for Agents Instead of Humans
  3. Anthropic's Playbook: From Five Workflow Patterns to Dynamic Workflows
  4. The Knowledge Graph as Shared Memory: Solving Context Bottlenecks
  5. Staged Build Path: From a Karpathy Loop to a Graph-Grounded Swarm
  6. Evaluation and Quality: Running Graph Autoresearch on Gold Sets
  7. When to Use (and When to Avoid) Knowledge Graphs in Agent Systems
  8. The Takeaway: The Architectural Shift to Graph Engineering
  9. References

When you move from isolated prompt execution to production autonomous systems, you face a fundamental bottleneck in context management and state representation. Graph engineering defines the architectural shift where autonomous agents transition from simple single-agent iteration loops into multi-agent systems grounded by explicit, durable state. Rather than relying on implicit context window memory that degrades across long runs, graph engineering structures execution history and domain facts as directed acyclic graphs (DAGs) and versioned knowledge graphs.

In a single-agent loop, an agent modifies code or parameters, measures execution against a fixed metric, and decides whether to keep or revert the change. As you scale to complex problems, this single-threaded model breaks down. The transition to multi-agent graph systems allows you to externalize iteration, parallel search, experiment lineage, and shared domain facts outside of individual chat transcripts. Each progressive pattern externalizes a distinct system bottleneck: a loop externalizes iteration and evaluation; a chain externalizes task order; a swarm externalizes parallel search and role specialization; a DAG externalizes experiment lineage; and a knowledge graph externalizes shared facts, provenance, and cross-session memory.

By shifting state into explicit graph structures, you enable asynchronous agent collaboration without polluting context windows. Instead of replaying full execution transcripts into orchestrator contexts, agents write structured updates to commit DAGs and knowledge graphs, enabling precise retrieval, provenance tracking, and verifiable evaluation across long-running swarms.

Graph Engineering architecture overview: Transitioning agent memory from volatile context windows into durable DAGs and Knowledge Graphs

Autoresearch and Karpathy's Self-Optimizing Loop

Andrej Karpathy's autoresearch model demonstrates the minimal executable harness for autonomous machine learning experimentation. Operating on a single GPU with roughly 630 lines of code, the harness ran approximately 700 experiments over two days, yielding 20 retained optimizations. These optimizations included QK normalization scaling, value-embedding regularization, AdamW parameter tuning, batchsize changes, depth changes, embedding learning rate, RoPE base frequency, targeted weight decay, initialization scale, and warmdown settings. The repository's accumulation of over 86,000 GitHub stars and 12,500 forks is not a scientific performance metric, but concrete evidence of the pattern's operational legibility and ease of reproduction across modern AI infrastructure setups.

Autoresearch execution harness and the 4 fundamental conditions for self-optimizing loops

The core system relies on three primary files that establish a strict execution boundary. prepare.py contains fixed data preparation and evaluation utilities that remain unmodified by the agent. train.py serves as the experimental surface where the model architecture, optimizer, hyperparameters, and training loop reside for the agent to edit. program.md provides natural-language instructions configuring the research process, constraints, evaluation metric, logging rules, crash handling, and autonomy policy. This setup embodies Software 3.0 context-as-interface programming, where natural-language prompts act as a programmable control interface configuring an entire autonomous research organization.

The loop succeeds because it satisfies four fundamental system conditions: verifiable output through a measurable validation metric, reversible action via git reset, a short execution horizon of approximately 5-minute runs, and a strictly bounded environment within the repository. During step 5 of the execution loop, the harness measures validation loss expressed as bits per byte (val_bpb) alongside peak GPU memory usage. If a trial crashes or fails to improve val_bpb, the system executes a reset; if val_bpb improves while respecting memory constraints, the commit is retained. You can implement this self-optimizing loop using a deterministic ratchet pattern:

python
def ratchet_loop(inspect, propose, apply, evaluate, keep, revert, better, baseline):
    history, current = [], baseline
    while True:
        state = inspect()
        change = propose(state)
        commit = apply(change)
        try:
            score = evaluate()
        except Exception as exc:
            revert(commit)
            history.append(Trial(commit, change, None, "crash", str(exc)))
            continue
        if better(score, current):
            keep(commit)
            current = score
            history.append(Trial(commit, change, score, "kept", ""))
        else:
            revert(commit)
            history.append(Trial(commit, change, score, "reverted", ""))

Autoresearch converts internal human working memory into machine-readable execution histories. Human researchers typically track hypotheses, parameter interactions, and code diffs in volatile cognitive state. The self-optimizing loop records every experiment alongside its parent commit, code diff, validation metric, and retain status, allowing autonomous agents to systematically explore search spaces, revisit prior lineages, and recover from failures without losing context.

From a Single Loop to AgentHub: Git for Agents Instead of Humans

Scaling from a single loop to multi-agent collaboration requires changing your version control assumptions. AgentHub inverts standard human Git workflows. Human development focuses on merging branches into a single canonical main branch through human-paced code reviews and pull requests. Agent research involves thousands of agents executing simultaneous experiments where most branches are never merged, failed attempts contain valuable search signal, and the primary operation is traversing a search graph rather than forcing branch convergence. AgentHub strips away convergence abstractions, operating without required main branches or merge queues.

AgentHub search DAG architecture and parallel exploration without branch convergence

AgentHub's minimal architecture consists of a Go server binary, a SQLite database, a bare Git repository on disk, an HTTP service, an API key per agent, a thin ah command-line interface, and a message board operating under strict rate limits and bundle-size constraints. However, because AgentHub was introduced as a conceptual sketch, productionizing this architecture requires you to solve critical engineering failure modes: distributed storage scaling, repository compaction, trust among untrusted agents, malicious code bundle execution, experiment reproducibility, semantic duplicate detection, compute scheduling, and long-term graph indexing.

You interact with the search space using dedicated CLI commands designed for graph operations:

bash
ah push          # Push HEAD commit to hub
ah fetch <hash>  # Fetch any commit
ah log [-agent X]# Show recent commits
ah children <hash> # Query ideas tried above commit
ah leaves        # Show unexplored search frontier
ah lineage <hash># Reconstruct ancestry path to root
ah diff <hash-a> <hash-b> # Compare any two commits

In AgentHub, the directed acyclic graph (DAG) is literally the system graph. Commit nodes store parent links directed to previous states, along with agent IDs, hypotheses, code diffs, validation metrics, execution runtimes, peak memory metrics, execution environments, keep-or-discard status, links to discussion posts, and links to related experiments. This graph structure allows agents to collaborate asynchronously across parallel search paths without copying full chat transcripts or replaying raw execution histories into context windows.

Anthropic's Playbook: From Five Workflow Patterns to Dynamic Workflows

Anthropic's architectural evolution moves from static composition patterns to programmatic dynamic orchestration. The 2024 foundational patterns—Prompt Chaining, Routing, Parallelization, Orchestrator-Workers, and Evaluator-Optimizer—rely on explicit developer-defined control flow. In a graph-grounded architecture, the graph fulfills distinct operational roles across these patterns: serving as a retrieval source in Augmented LLMs, a gate signal in Prompt Chaining, a classifier input in Routing, a shared surface in Parallelization, shared memory in Orchestrator-Workers, and an evidence grounding layer in Evaluator-Optimizer loops.

Anthropic Dynamic Workflows runtime in Bun and the 4-step Knowledge Graph construction pipeline

The 2026 Dynamic Workflows architecture shifts orchestration into generated JavaScript programs where the model dynamically writes workflow code at runtime. Triggered via the keyword "workflow" or "ultracode mode", Claude spawns up to 1,000 parallel sub-agents in fresh, isolated context windows, maintaining intermediate state directly in script variables. Operating with a default concurrency of 16 sub-agents, this model enabled real-world scale achievements, such as Bun porting approximately 750,000 lines of Zig to Rust in 11 days with 99.8% of tests passing.

javascript
const files = await tools.glob("src/**/*.ts");
const audits = await gather(
  files.map((file) =>
    spawn("auditor", { file, instructions: "Inspect for race conditions. Return JSON." }),
  ),
  { concurrency: 16 },
);
const suspicious = audits.filter((r) => r.confidence >= 0.7);
const reviews = await gather(
  suspicious.map((r) =>
    spawn("reviewer", { report: r, instructions: "Try to refute this finding." }),
  ),
  { concurrency: 16 },
);
return await spawn("synthesizer", { audits, reviews, instructions: "Produce one cited report." });

To support structured knowledge representation, Anthropic's Knowledge Graph Construction Cookbook defines a four-step pipeline that replaces classical NLP pipelines with structured model prompts: Extraction (using Claude Haiku to parse text into typed entities and relations), Resolution (using Claude Sonnet to cluster entity candidates), Assembly (building a NetworkX MultiDiGraph with provenance), and Querying (serializing subgraphs for Sonnet reasoning with citations).

python
class Entity(BaseModel):
    name: str
    type: EntityType
    description: str
 
class Relation(BaseModel):
    source: str
    predicate: str
    target: str
 
def extract(text, client) -> ExtractedGraph:
    response = client.messages.parse(
        model="claude-haiku-4-5",
        messages=[{"role": "user", "content": PROMPT.format(text=text)}],
        output_format=ExtractedGraph,
    )
    return response.parsed_output

Entity resolution transforms raw surface forms into canonical graph nodes using contextual reasoning rather than simple string matching. By grouping candidates by entity type and evaluating detailed descriptions, models resolve non-overlapping surface variations such as "Edwin Aldrin" and "Buzz Aldrin". Resolution operations must remain additive, inspectable, and reversible, preserving original aliases, source documents, resolution rationale, confidence scores, and run IDs so incorrect merges can be rolled back without re-running document pipelines.

The Knowledge Graph as Shared Memory: Solving Context Bottlenecks

A knowledge graph acts as a central shared memory layer across multi-agent architectures, addressing three critical system roles. As Shared Memory, worker agents publish structured graph updates rather than dumping raw transcripts into an orchestrator's context window. As a Grounding Layer, evaluator agents verify generated claims against explicitly linked edges in the graph. As a Persistent World Model, the graph preserves cross-session state, active claims, temporal relationships, and audit trails long after individual model context windows are reset.

Knowledge Graph as a shared memory layer and Cross-Graph Linkage between Commit DAG and facts

Worker agents interact with this memory layer through atomic, schema-validated graph updates:

python
@dataclass
class GraphUpdate:
    nodes: list[dict]
    edges: list[dict]
    run_id: str
    agent_id: str
 
def publish(update, graph, validator):
    validator.check_schema(update.nodes, update.edges)
    validator.check_provenance(update.nodes, update.edges, update.run_id)
    with graph.transaction() as tx:
        tx.upsert_versioned_nodes(update.nodes)
        tx.add_edges(update.edges)
        tx.link_run(update.run_id, update.agent_id, update.nodes, update.edges)
        tx.commit()

You must enforce a strict structural boundary between the Commit DAG and the Knowledge Graph while maintaining explicit logical traceability between them. The Commit DAG tracks work lineage, code diffs, and experiment iterations (answering what changed and which agent modified it). The Knowledge Graph tracks domain entities, claims, source documents, and real-world facts (answering what is known and what supports it). In production platforms, you connect them via concrete edge syntax:

text
(agent_run_183) -produced-> (claim_441) -modified-> (commit_a81f) -evaluated_by-> (evaluation_92)
(claim_441) -about-> (entity_autoresearch) -supported_by-> (source_readme) -supersedes-> (claim_238)

Every graph write in your infrastructure must satisfy four system invariants: (1) every claim has a source or is explicitly marked as inference; (2) every artifact identifies an authoring run and version; (3) every evaluation identifies a concrete rubric; and (4) every superseded object remains addressable. To construct execution context without context dumping, you must retrieve task-specific subgraphs by resolving task entities, expanding 1–2 hops over allowed edge types, including current artifact versions, prioritizing recent verified claims, explicitly including active contradictions, and serializing within strict token budgets using stable edge identifiers for precise citation.

Staged Build Path: From a Karpathy Loop to a Graph-Grounded Swarm

Building graph-grounded multi-agent systems requires an incremental implementation path to manage system complexity. On Day 1, you build a basic reflective single-agent loop (reflective_task) that generates drafts, evaluates them against explicit rubrics, and revises until criteria are met:

python
def reflective_task(task, gen, eval, max_rounds=3):
    versions = [gen(task)]
    for _ in range(max_rounds):
        review = eval(task, versions[-1])
        if review["decision"] == "approve":
            return {"result": versions[-1], "versions": versions}
        versions.append(gen(task, prior=versions[-1], instructions=review["changes"]))
    return {"result": versions[-1], "status": "iteration_limit"}

On Day 2, you integrate typed tools with strict schema validation and execution permissions. By Week 1, you introduce dynamic JSON planning only when task paths vary, validating task dependencies before execution using a strict template:

json
{
  "objective": "Construct verified knowledge graph",
  "steps": [
    {
      "id": "s1",
      "action": "extract",
      "input": "documents",
      "depends_on": [],
      "success": "Each doc returns ExtractedGraph"
    },
    {
      "id": "s2",
      "action": "resolve_entities",
      "input": "s1.entities",
      "depends_on": ["s1"],
      "success": "Every surface form maps to one ID"
    },
    {
      "id": "s3",
      "action": "assemble_graph",
      "input": ["s1.relations", "s2.mapping"],
      "depends_on": ["s1", "s2"],
      "success": "All endpoints resolve to nodes"
    }
  ]
}

By Week 2, you implement multi-agent separation of concerns across six specialized roles—planner, implementer, test author, reviewer, security reviewer, and synthesizer—enforcing artifact contracts at handoffs and using Git worktree isolation when multiple agents modify the codebase. By Month 1, you wire operations into a versioned persistent graph storing entities, claims, sources, relations, artifacts, agent runs, evaluations, and versions. By Month 2, you scale execution to swarms handling embarrassingly parallel workloads, enforcing strict reducers, worker concurrency limits, total worker caps, per-worker timeouts, retry policies, token budgets, and final evaluation gates.

Staged build path from reflective loop to swarm and the 5-plane reference production architecture

You must isolate system capabilities across a five-plane reference production architecture: the Control Plane (manages objectives, plans, and budget allocation), the Execution Plane (runs tools, sub-agents, and code in isolated environments), the Artifact Plane (stores immutable plans, code diffs, reports, and evaluations), the Graph Plane (manages entities, claims, commit DAGs, and provenance), and the Evaluation Plane (executes deterministic tests, model rubrics, and statistical scorers).

Evaluation and Quality: Running Graph Autoresearch on Gold Sets

You maintain graph system quality by applying Karpathy's ratchet loop directly to graph infrastructure—a methodology termed Graph Autoresearch. Instead of optimizing model weights or hyperparameter configurations, the ratchet optimizes extraction prompts, ontology definitions, entity resolution policies, and context serialization logic against an established benchmark gold set.

You must measure system health across distinct architectural layers while avoiding common metric misreadings:

  • Extraction: Measure Entity/Relation F1, precision, recall, and schema validation rate. Misreading: High extraction precision frequently masks low recall and missing entities.
  • Resolution: Track compression ratio (raw surface forms divided by canonical entities), pairwise precision/recall, false merge rate, and missed merge rate. Misreading: High compression ratio is not inherently positive; aggressive over-merging destroys graph accuracy by creating false connections.
  • Graph Topology: Monitor component count and graph density. Misreading: Collapsing a graph into a single connected component indicates catastrophic over-merging rather than optimal synthesis.
  • Query & Workflow: Validate accuracy, multi-hop path validity, citation correctness, and task success versus token cost. Misreading: Fluent model answers can cite irrelevant edges, and adding sub-agents can increase overall activity without adding signal.
  • Operations: Monitor subgraph latency, stale entity count, graph update failure rates, and agent retry rates.

Production monitoring requires tracking graph topology metrics continuously over time. Sudden spikes in disconnected components signal entity resolution regressions where surface forms fail to merge. Conversely, sudden drops in component counts signal aggressive over-merging that collapses distinct domain entities into invalid central hubs.

To stress-test graph stability, you must construct adversarial gold set test cases containing misleading aliases, contradictory event dates, and disconnected execution paths. When evaluators detect missing or contradictory edges during claim verification, they return structured failure feedback outlining missing source paths, allowing agents to execute targeted retrieval rather than unguided context regeneration.

When to Use (and When to Avoid) Knowledge Graphs in Agent Systems

To evaluate whether your architecture requires a knowledge graph, answer six selection questions: 1) Can success be verified deterministically or via explicit rubrics? 2) Are execution steps stable over time? 3) Are subtasks independent and parallelizable? 4) Must alternative search lineages remain accessible in a DAG? 5) Must facts and state survive individual agent runs? 6) Can the organization afford the cost and latency overhead?

Every graph-grounded agent run must declare explicit complexity budgets prior to execution. You must enforce hard ceilings on total model calls, concurrent sub-agents, max concurrent workers, tool calls, wall-clock execution time, token consumption, financial expenditure, retry attempts, total graph writes, and minimum evidence required for finalization. When a budget limit is reached, the system must return its best current artifact, completed work, and unresolved issues rather than continuing unbounded execution.

You must explicitly avoid introducing a knowledge graph when simpler architectural patterns suffice. Knowledge graphs add unnecessary complexity when tasks are independent, context is restricted to single documents, relations are simple and static, standard relational databases fulfill query requirements, or extraction noise exceeds graph traversal value.

Graphs earn their operational overhead only when multi-hop queries, persistent cross-session state, evolving domain relationships, and verifiable provenance are essential to system success. By matching architectural complexity directly to verified task requirements, you prevent resource exhaustion while maintaining system transparency.

The Takeaway: The Architectural Shift to Graph Engineering

The transition from vibe coding (implicit model outputs) to agentic engineering (orchestrated chains and tools) culminates in graph engineering (explicit, durable, shared memory). Performance bottlenecks in advanced agent systems are rarely caused by model intelligence limits, but rather by poor memory placement and inadequate evaluation mechanics.

By externalizing state into explicit commit DAGs and versioned knowledge graphs, you establish systems with clear provenance, verifiable execution histories, and persistent domain facts. Every critical output in a production agent system must be traceable to an objective, a plan, an artifact, a source, a graph path, an evaluator decision, and a bounded execution record.

References

Share this article