There is a three-way split going around for describing agent systems, and it is a good one. The harness is the environment: tools, memory, permissions, what goes into context. The loop is the cycle: execute, check the result, decide whether to go again, stop. The graph is the path: which step follows which, where it branches, what runs in parallel, where a human signs off.
Clean definitions. The usual problem with clean definitions is that they get memorised and change nothing. A taxonomy earns its keep only if it makes a decision easier, so here is the decision it makes easier.
Your agent just did something wrong. What do you edit?
The default answer is the prompt, and the prompt is not in any of the three layers. That is not an accident of the framing — it is the point of it.
Symptoms sort cleanly
Most agent misbehaviour announces its layer, once you know what to listen for.
| What you see | Layer | Why |
|---|---|---|
| Calls a tool that does not fit the request | Harness | The candidate set or its descriptions are wrong |
| Confidently states something stale or false | Harness | What is in context is wrong or missing |
| Tries, fails, tries the identical thing again | Loop | No evidence is entering the retry decision |
| Runs until you kill it | Loop | No stop condition, or one that cannot trigger |
| Produces something almost right, forever | Loop | Nothing distinguishes good enough from not |
| Does the right steps in the wrong order | Graph | The path is implicit and the model is guessing it |
| No handling for a case you knew about | Graph | The branch was never drawn |
| Acted where a human should have approved | Graph | No gate on that edge |
Two things fall out of this table immediately.
The first is that “the model is not smart enough” appears nowhere. Every row above is a property of the system built around the model, and every row has a fix that is not a bigger model.
The second is that the fixes are unrelated to each other. Adding a tool does nothing for a loop with no stop condition. A stop condition does nothing for a path with a missing branch. Which is why “improve the prompt” — the one intervention people reach for regardless of symptom — resolves roughly a third of cases and gets applied to all of them.
The misdiagnosis that costs money
One confusion is far more common and far more expensive than the rest: a harness problem treated as a loop problem.
It looks like this. The agent fails a task. Somebody adds retries, maybe an escalation to a stronger model on the second attempt. The failure rate improves a little, because occasionally sampling gets lucky, so the fix looks validated.
What actually happened is that the agent lacked a tool, or lacked a piece of context, or was choosing between two tool descriptions that read alike. None of that changes between attempt one and attempt four. You have taken a failure that cost one model call and made it cost four, plus latency, and you have hidden the diagnosis under a retry count.
The tell is simple and worth checking before adding any retry logic: do the failed attempts differ from each other? If attempt three is doing something genuinely different from attempt one, the loop is working and retries are earning their cost. If the attempts are near-identical, nothing new is entering the decision, and no number of them will help. That is a harness gap, and it usually resolves to a missing tool, missing context, or two tools the model cannot tell apart.
The mirror-image mistake exists too and is cheaper: a loop problem treated as a harness problem. The agent never converges, so somebody adds more tools and more context, and the extra tools make selection worse. That direction at least fails visibly.
Each layer has a different test
The most practical consequence of the split is that the three layers are verified in three genuinely different ways, and confusing the methods produces tests that pass while the system is broken.
Harness is testable statically. Does this agent have the tool it needs for this task class? Does the assembled context fit, and is what matters near the top rather than buried where attention thins out? Is anything stale in there? You can assert these before running anything, and they are cheap, deterministic assertions — the kind that belong in CI.
Loop is only testable as a distribution. One run tells you almost nothing, because the interesting quantities are rates: how often does it converge, in how many turns, how often does it hit the cap. A loop that succeeds nine times in ten and burns forty turns on the tenth is a specific, fixable defect, and no single execution reveals it. This is the layer where you need many runs and a number, not a pass.
Graph is testable as coverage. Every node reachable, every branch exercised, every terminal state actually terminal. This is ordinary path coverage and the classic bug is boring: an error branch that was drawn but never executed, so nobody noticed it points at a node that does not exist.
Three layers, three test strategies, and one combination reports green while broken: solid harness assertions, no loop distribution at all, and an error branch nobody ever executed.
Where the boundaries actually blur
Being honest about the framing: the split is cleaner in a diagram than in a codebase.
A retry with escalation to a stronger model is a loop decision that changes the harness. A branch that adds a tool for one path is graph structure altering harness contents. And the stop condition — the single most important thing in the loop layer — is usually a check against something the harness supplied, which means a bad harness produces a stop condition that fires on the wrong evidence.
That last one is worth stating plainly, because it defeats the tidy separation: a loop can only be as good as the signal it stops on. If your convergence check is a model judging its own output, the loop inherits that judge’s accuracy, and a loop that stops early on a judge that is wrong a third of the time is not a loop problem at all.
So use the layers to localise, not to partition. The question “which layer” is a first move that narrows the search, not a claim that the fix lives entirely in one place.
The order to check
When something goes wrong, in this sequence:
- Did the attempts differ from each other? No means harness, and stop adding retries.
- Did it have what it needed? The right tool, current context, permission to act. Cheapest to check, most often the answer.
- Did it know when to stop, and on what evidence? If the stop condition is a judgement, its accuracy is now the loop’s accuracy.
- Was the path drawn, or inferred? Steps in the wrong order and missing error handling both mean the sequence lived in the model’s head rather than in your structure.
- Only then, the prompt. It is genuinely the fix sometimes. It is the first thing tried almost always, which is why so much prompt tuning produces so little.
The value of the three-layer split is not that it describes agent systems accurately, though it does. It is that it gives you four questions to ask before the one everybody asks first.