The setup is straightforward enough to describe in a sentence: connect an agent to your issue tracker and your repository, point it at a ticket, and get a pull request with tests in it.
The wiring is the easy half and it is what every writeup covers. The half that decides whether this is useful or a novelty is the instruction file — the document that tells the agent how this specific repository behaves. And the surprising thing about the instruction files that actually work is what they spend their space on.
They are mostly telling the agent not to improve things.
Look at what the useful rules say
Here is a representative set, lightly trimmed, from a repository where this pipeline runs for real:
- Preferred style for new code: `const`, `ReadonlyArray<T>`, spread operators.
When editing existing test code, follow the local file pattern if it
already uses `let` or mutable builder state.
- Prefer specific types over `any`, but preserve existing `any` usage in
legacy test files unless the change is already refactoring that code.
- Match the existing import style of the file being edited. Do not rewrite
established relative imports just to force absolute paths.
- File names kebab-case, classes PascalCase, variables camelCase — unless
an existing file already follows a different established pattern.
Every one of those rules has the same shape: here is the preference, and here is when to ignore it. Three of the four exist purely to stop the agent from tidying code it was not asked to touch.
That is not what people put in these files. The usual instruction file is an aspirational style guide — the code we wish we wrote, stated as rules. Which reads as obviously correct and produces a specific, predictable disaster.
Why the style guide backfires
A human given a style guide applies judgement about scope. They are fixing a bug in one function, they notice the file around it uses an older pattern, and they leave it alone — because they know a reviewer wants to see the fix, not the fix plus two hundred lines of unrelated normalisation.
An agent given the same style guide applies it. Every file it opens gets brought into compliance, because nothing told it that compliance has a boundary, and because “make the code match the standards” is a perfectly reasonable reading of what you wrote.
The output is a pull request where eight lines are the feature and two hundred are churn. And that diff is not merely annoying, it is unreviewable in the specific way that defeats the entire pipeline:
- The reviewer cannot see the change, so they skim.
- Skimming a large diff means the eight lines that matter get the same attention as the two hundred that do not.
- Every touched file is now a merge conflict for somebody else’s branch.
- If something breaks next week,
git blamepoints at a formatting sweep.
The whole arrangement — ticket in, PR out — rests on a human reviewing the result. A change that makes review harder does not trade off against the benefit; it removes it.
So the rules that constrain scope are not stylistic preferences. They are what keeps the artefact reviewable, which is what keeps the pipeline honest.
The other half: mechanics an agent cannot guess
The second category of rules that earns its space is the one describing how the repository physically works.
1. Create `test/tests/<feature-name>.ts`.
2. Export a single named function matching the feature.
3. Add the export to `test/tests/index.ts`.
4. Call it inside `test/my-test-1.ts` under the RUN_PERFORMANCE_TEST guard.
Nothing here is about quality. It is where files go, what to export, and — the important one — the registration step. Step three is invisible from any single file. An agent can read your entire test directory and write a perfect test, and if it does not add the line to index.ts the test never runs.
That failure is the worst kind, because it is silent. The suite is green. The PR looks complete. Nobody notices for a month.
Registration steps, guard flags, generated files that must be regenerated, the one config that must be updated in two places — these are the highest-value lines in an instruction file, and they are the ones people skip, because everyone on the team has them memorised and they do not feel like rules.
The test for whether something belongs here: would a competent new hire get this wrong on their first PR? If yes, write it down. An agent is a new hire on every single task, with no memory of the last one and no ability to ask.
What not to put in the file
Given limited attention, three things are worth omitting.
Anything the linter enforces. If the formatter reformats and CI rejects non-compliance, the rule is already enforced by a mechanism that does not consume context. Repeating it costs tokens and buys nothing.
Anything derivable from reading a nearby file. “Tests use describe and it” is visible in every test. The agent reads the neighbourhood. Spend the space on what the neighbourhood does not show.
Philosophy. “Write clean, maintainable code” changes no output. It is the same failure as a vague acceptance criterion in a ticket — it feels like a requirement and it constrains nothing, which is the difference between a specification and a wish.
The review still decides everything
One more thing worth stating plainly, since these writeups tend to end on the automation and not on the gate.
The pipeline produces a pull request, not a merge, and that distinction is load-bearing. Everything above is in service of one property: that a human can look at the result and tell in a couple of minutes whether it is right.
Which suggests the metric to watch is not how many tickets the agent closed. It is the size of the diff relative to the size of the change, and the fraction of PRs a reviewer approves without asking for scope reduction. When those go bad, the instruction file is missing a restraint, and you can usually name which one by looking at what the agent tidied.
Same principle as any other automation you would put in front of a person: the output is only as valuable as it is checkable, and a large output is not more valuable, it is less checkable.
The ticket is the other input, and it is usually worse
Everything above is about the repository side. The other input is the ticket, and it gets almost no attention in these writeups because it belongs to a different team.
A ticket written for a human is a prompt for a conversation. “Add coverage for the new filter behaviour” works fine when the person reading it sits near whoever wrote it and can ask which filter, on which endpoint, and what counts as correct. An agent reading the same sentence produces a plausible test for a behaviour nobody requested, and it produces it confidently.
The failure is quiet in the same way the missing registration line is quiet. The PR arrives, the tests are green, the code follows every convention in the instruction file, and the thing being asserted is not the thing that needed asserting. A reviewer skimming for style compliance will pass it.
Two practical consequences follow.
The tickets that work are the ones with a concrete expected result in them. Not more words, more specifics: which endpoint, which inputs, what the response should contain. A ticket that names an exact expected value is a ticket an agent can satisfy verifiably; a ticket that describes an intention is not.
Route the vague ones to a human. The pipeline does not need to accept every ticket, and the ones it should refuse are recognisable before any work happens: no acceptance criteria, no named endpoint, no example. That check costs nothing and it removes the failure mode that review is worst at catching.
Where to start
If you are writing one of these files today, the order that gets results fastest:
- Write down the registration and wiring steps first. The things that fail silently. These pay off immediately.
- Write the restraints second. For every style preference, add the sentence that says when to leave existing code alone. “Follow the local pattern” is the single most useful phrase available.
- Delete anything the linter already enforces. Free context.
- Watch the first ten PRs and add a rule per surprise. The file should grow from observed failures, not from imagining what the agent might get wrong.
Point four matters most in the long run. Instruction files written up front encode what you assumed; instruction files grown from real PRs encode what actually went wrong, and those are rarely the same list.