Your Agent Is a While Loop: Harness, Loop, and Graph Engineering
Human thoughts, AI-assisted write-up.
Strip an agent back to its skeleton and it's embarrassingly simple. A while loop.
while not done:
context = observe() # read files, tools, memory, the last result
thought = reason(context) # what should I do next?
result = act(thought) # run the command, edit the file, call the API
done = check(result) # did it actually work?
That's it. Everything else is production hardening around those five lines. A better editor integration, a permissions model, a memory layer, checkpoints, retries, a slicker approval prompt. More tools, more hardening, more DevX. Useful, all of it, but none of it changes the shape.

The real problem here is a naming one.
Every few weeks the timeline discovers a new "-engineering". Prompt engineering, then context engineering, now loop and graph engineering, with eval engineering warming up in the bullpen. Some of it's real. A lot of it is old ideas in a new jacket (and "graph engineering" has picked up an official ring on Twitter it hasn't earned, which I'll get to), sold as a discipline you're somehow already behind on.
You aren't. The moment you open a coding agent you're in a loop: a ReAct-style read, think, act cycle is loop engineering, you just didn't call it that. The RALPH crowd made "wrap the agent in a while loop and let it run" a named technique, and then it quietly went native in Claude Code, Codex and the rest. You were doing loop engineering the first time you let an agent retry a failing test.
Under the hype there's one idea worth keeping: reliability lives in three layers, and mixing them up is how you burn a week "prompt engineering" a problem that was never in the prompt.
- The harness is the environment the loop runs in.
- The loop is the feedback that decides whether to go again.
- The graph is the flow, once the work grows branches.
I learned that the expensive way in July, when three frontier models chased the same security scare on my desk and the cheapest one landed in the same price band as the dear ones. I'll come back to that. First, which layer breaks, and how you tell.
The stack, and which clock each layer fails on
| Layer | Unit | Question it answers | Failure clock |
|---|---|---|---|
| Prompt engineering | one message | what do I tell the model? | seconds |
| Context engineering | one window | what goes in and out? | one answer |
| Harness engineering | one run | what can the model do, safely? | one run |
| Loop engineering | repeated runs | how does the system check, retry, stop? | hours to days |
| Graph engineering | whole workflow | what path can the work take? | the whole process |
Most agent failures get blamed on the model. Usually the wrong layer.
- Your harness is broken if the agent can't call the right tool, loses state between runs, leaks a permission, or can't be inspected. The model isn't your core problem.
- Your loop is broken if the agent almost works, but keeps stopping early, retrying blindly, or declaring success with no evidence.
- Your graph is broken if the process has branching, handoffs, approvals, parallel work and recovery paths, and none of it is visible anywhere.
Harness: the floor the agent stands on
Ever watched an agent report a green deploy off a curl that came back 200 with a stack trace in the body? That one does my head in every time. The API swallowed the error, the agent trusted the status code, and no model on earth was going to reason its way out of that. It's a harness bug.
The harness is everything that decides whether the model can operate at all: tools, memory, permissions, isolation, and some way to see what actually happened. Two teams can run the identical model and get wildly different outcomes, because one gives clean tool schemas, stable state and traceable execution, and the other gives a vague prompt and nothing to verify against.
The smell is nearly always one of three: a tool it was never given, state that doesn't survive a restart, or an error swallowed on the way back. Fix those and the model looks cleverer overnight. A smarter model doesn't grow a tool it was never given.
Loop: feedback, not vibes
The loop is the controller wrapped around the model. Trigger, work, gate, state, stop. The last three are where the reliability lives: a gate a machine can check, state that outlives the conversation, and a stop with a hard cap on it.
The one rule that matters:
Don't loop on confidence. Loop on evidence.
"The agent says it's done" is a hope, not a stop condition. Real stop conditions are things a machine can check without asking the model's opinion: tests pass, schema validates, citations resolve, CI goes green, the reviewer approves. Prompting improves one response. A loop improves the process that produces responses.
And the trap inside every loop is the evaluator. Who checks the checker? The generator should never be the sole judge of its own work, because the same context that produced the work also contains all the arguments for why the work is fine. Self-review compounds blind spots. A real evaluator looks nothing like the thing it's checking:
- Fresh context, so it doesn't inherit the generator's rationalisations.
- Separate instructions, ideally a different model.
- It acts on the output instead of re-reading it: opens the page, clicks the button, inspects the DOM, screenshots the result.
- Its default stance is assume broken until the evidence proves otherwise.
A UI evaluator that says "the JSX looks fine" isn't an evaluator. It's the generator wearing a hat.
Graph: the shape of the work
A prompt is a sentence. A loop is a cycle. The graph is the shape of the work itself: what can run now, what has to wait, what runs in parallel, what branches, where a human gate sits, where it can resume.
The primitives are boringly familiar to anyone who has drawn a DAG. If you've written Terraform, you've already described one: resources that depend on other resources, a plan that works out what can run in parallel and what has to wait. Same shape turns up in a CI pipeline, or in LangGraph back in early 2024, when it was the tool everyone reached for to build the first non-conversational agents. Nodes, edges, a bit of state:
| Primitive | What it is |
|---|---|
| Node | a bounded job: an agent call, a deterministic function, a human review, a tool run |
| Edge | a dependency, a data contract between two nodes |
| State | durable context carried across hops |
| Router | a conditional branch based on validated output |
| Fan-in | the next step needs the whole upstream set before it can run |
| Cycle | an explicit loop edge with a convergence rule |
| Verifier | a node whose job is to try to kill a finding before anything downstream trusts it |
One question sorts a real edge from a fake one, asked at every "and then": does the next step read the previous step's output? If not, you invented a dependency, and you pay for it in wall-clock. A straight line A to B to C to D is the degenerate graph: sometimes right, often slow for no reason.
Most real work is neither a straight line nor anything exotic. Take shipping a change. The checks don't depend on each other, so they fan out and run at once. The merge depends on all of them, so it waits. And look at the two dashed edges: getting to all green is itself a loop (red, fix, push, re-run, until the gate passes), and so is review. Two loops sitting inside one DAG. That's a graph with cycles in it, and you've drawn it a hundred times without calling it one:
A few habits carry most of the weight:
- Cut invented edges. Independent steps run in parallel, not one behind another.
- Reduce in code, not in a model. Dedup, filter, rank, join are deterministic; a hash set and a sort are exact, free, instant. Handing forty findings to a model to "remove the duplicates" is fuzzy, costs tokens, comes back different every run. Save the model for the judgment step.
- Judgment at the node, routing in code. The model can classify a result; code should pick the path. And plenty of nodes shouldn't be a model at all. Don't burn tokens asking an agent to
/security-reviewa diff when a scanner will do it deterministically: Semgrep, CodeQL or gitleaks for the SAST pass, OWASP ZAP for DAST against the preview deploy, and there's one for every language (Bandit, gosec, Brakeman). Feed what they find into a fix-and-rescan loop, and the model pays for the judgment, not the scanning. - Loop until dry for unknown-size work. Keep spawning finders until a couple of rounds turn up nothing new, and dedupe against everything you've ever seen, rejected candidates included, or the loop pays to rediscover the same dead end every round.
Mind you, most of the time a straight line is fine. If the job is three steps and forty seconds, drawing a graph for it is faffing about. The graph earns its keep when the wall-clock or the token bill says so, and not before.
A bit of misinformation worth clearing up: "graph agents" aren't GraphRAG
Two completely different things wear the word "graph", and they keep getting mashed together. We've seen this film before: when GraphQL landed, half the internet treated it as a graph database, because both had "graph" in the name. They have nothing to do with each other. GraphQL is a query language for an API, a graph database is a storage engine. Same word, different room.
Same thing here. Graph engineering is the execution graph, the path your work takes. GraphRAG is a memory retrieval mechanism: it stores your documents as a knowledge graph so the model has something to fetch from.
| "graph" in graph engineering | "graph" in GraphRAG | |
|---|---|---|
| A node is | a job, step or agent | an entity |
| An edge is | a data contract | a relationship |
| The graph is | the execution path | the knowledge index |
| It decides | how the work runs | what gets retrieved |
| Examples | LangGraph, AutoGen | Microsoft GraphRAG (2024) |
One line: storing your knowledge in a graph doesn't make your agent's control flow a graph.
Claude's primitives, and what each one is
In Claude Code the layers have names, and the mistake I see most is confusing a trigger with a stop condition: /loop decides when to fire again, /goal decides when you're done.
| Reach for | When |
|---|---|
/loop |
A prompt should fire again on an interval or self-paced, and you decide when to stop |
/goal |
You have a clear done-condition and want it to run across turns until that's true |
/schedule |
The work should run on cron in the cloud with no session open |
| workflow / ultracode | The task has real shape: parallel branches, routing, human gates. Ultracode lets Claude draw that graph for you |
| saved workflow | You want that shape versioned and re-runnable |
None of this is unique to Claude Code. Codex has background cloud tasks and a scriptable codex exec; GitHub Copilot's coding agent triggers off an assigned issue and runs to a PR. Different names, same two organs every harness grows: something that triggers the work, something that decides it's done.
Subagents versus workflows
A subagent and a workflow aren't the same lever, and reaching for the wrong one is the difference between a fan-out that holds and one that quietly loses work.

- Subagent: the plan lives in the model's head. A fresh agent with a clean context that hands back one answer. Great for a one-off job or a few parallel reads. But the main agent re-decides the plan every turn, every result piles back into its context, and nothing guarantees every branch finished or got checked.
- Workflow: the plan lives in code. Claude writes the script once. After that the loops, routing, dedupe and fan-in run as plain JavaScript, and the model only fires at the
agent()calls. Same shape every run, room for hundreds of agents, and the checks are built into the flow instead of left to memory. - Rule of thumb: one job, send a subagent. A process you'll run twice, write a workflow.
Anthropic's dynamic workflows cookbook draws the same line: "the plan lives in code instead of context". OpenProse got there from another direction, with contracts that wire steps by what they require and ensure, and I still keep it on hand.
Teammates: when the agents talk to each other
Subagents and workflows both keep one orchestrator in charge. Agent teams drop that. Switch them on (CLAUDE_CODE_EXPERIMENTAL_AGENT_TEAMS=1, still experimental) and each teammate is a full Claude Code session in its own tmux pane, talking to the others by name instead of routing everything back through you. It's heavier than a fan-out, so save it for when the agents need to argue: competing hypotheses on a stubborn bug, or a security lens and a correctness lens that should disagree out loud rather than file in parallel.
godmode: what happens when you compose all of them
Every Thursday evening I get the same low-grade anxiety. A couple of my plans reset on Friday evening, and whatever tokens I haven't burned by then just vanish: paid for, never used. So the question stopped being which model to use and became how to put that spare capacity to work while I'm not at the keyboard. godmode is what I came up with.
Advisory: this is a token guzzler. It's there to build things purely autonomously, without you needing an end game. Point it at capacity you'd lose anyway, not at a budget you care about.

Point it at a north-star outcome and it runs as a perpetual product factory, inventing the next epic and shipping it while it regresses everything so far. It's the clearest thing I've built for showing the three layers stacking:
- Harness. Hooks enforce what the model would skip: every state write publishes the dashboard, a Stop hook won't end the turn without the phase explainer, and a lease stops two machines driving at once.
- Loop. A driver re-arms itself every ten minutes and never marks itself DONE: regression runs alongside each new generation, a real defect jumps the queue, and only budget, deadline, safety or a lost lease stop it.
- Graph. Every stage is a workflow with the spec inside it, and every creative call goes to two models with the dissent kept. Research is the exception: one model, judged by its citations, writing to a ledger it rereads every generation, which is that dedupe-against-everything rule from earlier turned on its own memory.
- Finite or perpetual. Same machine, one flag.
There's no new magic in it. Harness plus loop plus graph, turned up until the interesting engineering is entirely in the control surface, not the model.
A few things it's built that I actively use:
- sayso, push-to-talk dictation for Android and macOS.
- cerebro, a local-first, token-minimal daily tech-intelligence pipeline.
- agent-bridge, which lets agents drive, inspect and debug a real browser.
Plus a bunch I abandoned, which is fine when the tokens would have evaporated anyway.
Four shapes, from real work
A loop and a graph aren't rivals. A graph can hold loops; a loop can be one cycle inside a graph. The point is to see which shape the work already has. Two of the everyday ones, side by side:
- CI fixer, a loop. Trigger on a failed run, read the logs, patch, gate on test plus lint plus build, journal each attempt as state, stop on green or after three tries with an escalation. One failure, one fix path, no graph.
- PR review, a graph. Fan out across lenses, each its own reviewer node; put a verifier in front of every high-severity claim (the fresh-context evaluator from the loop section, placed as a node); fan in to dedupe and rank; synthesize. The win is the confidence structure around the reviewers, not more of them.
- Deep research, a graph. Scope, search sources in parallel, fetch, dedupe in code, adversarially verify claims, synthesize, gate on citations resolving, hand to a human. The topology is the product.
- Scheduled digest, a trigger around a graph. Fire daily with
/schedule, fan out across releases, changelogs, docs, reduce and rank by impact, synthesize, keep loop state so you never report the same thing twice.
Where these go wrong is hidden topology, almost never the model: a serial chain where the steps were independent, a model burning tokens on a deterministic transform, a self-review loop approving itself, a retry loop with no cap, a dedup that only checks confirmed findings so the dead ends keep coming back.
Back to that July scare. A security alarm on my desk turned into an accidental benchmark: I pointed Fable, GPT-5.6 Sol and Kimi K3 at the same question, and all three landed on the same answer. Not a compromise, an Apple Continuity handoff.
Then the bill came in. Fable took about 27 minutes. K3, cheapest per token by a mile, took 51 and chewed through roughly 2.9 million tokens against Fable's 650 thousand, so the bargain landed in the same price band by the time the answer arrived.
Same answer, same price, nearly twice the wait. The extra spend bought more trips around the loop: more re-reads, more retries, more of the same context hauled back in. I wrote it up as measure cost by outcomes, and it's the lesson this whole post keeps circling. When a run costs too much, look at the loop and the graph before you blame the model.