Harness Engineering, Explained: What Makes AI Coding Agents Reliable (Learn Harness Engineering and MemoHarness)
Harness engineering is everything you build around a model so an agent finishes real work: instructions, tools, state, verification and scope. What Anthropic measured, what the Learn Harness Engineering course and the MemoHarness paper add, the gotchas, and how an FDE sets one up at a customer.
BY LUKAS HOFFMANN · FDEINTERVIEWS EDITORIAL · UPDATED OCTOBER 7, 2026 · 10 MIN READ
PRACTICE THIS:RAG and agent design questions ·Agentic evals and trajectories, the concept ·The Production AI Agents course ·The must-know FDE questions
Harness engineering is the work of building everything around a language model so that an agent finishes real tasks reliably: the instructions it reads, the tools it may call, the state it keeps on disk between sessions, the checks that decide whether work is done, and the limits on what it may touch. The model decides what code to write. The harness decides when, where and how, and whether the result counts. The term spread through late 2025 and 2026 as OpenAI and Anthropic published how they run coding agents for hours at a time, and two open-source projects now teach and test it: the Learn Harness Engineering course from walkinglabs (MIT, about 19,400 GitHub stars as of 7 October 2026) and the MemoHarness research code, which tries to learn harness settings from the agent's own runs. This post explains the pattern, what each project adds, the gotchas, and how a forward deployed engineer sets a harness up inside a customer's repository.
Same model, different harness: the evidence
The clearest data point comes from Anthropic's engineering post on harness design for long-running application development (24 March 2026). The author gave the same model the same one-sentence prompt, a 2D retro game editor, twice. Run alone, it worked for about 20 minutes, cost about 9 dollars, and produced an app whose play mode did not work. Run inside a three-agent harness (a planner, a generator and an evaluator), it worked for about 6 hours, cost about 200 dollars, and produced an editor where you could actually move a character and play a level.
| Solo agent | Planner, generator and evaluator harness | |
|---|---|---|
| Duration | about 20 minutes | about 6 hours |
| Cost | about $9 | about $200 |
| Scope | what the model chose to build | a planner expanded the prompt into a 16-feature spec over ten sprints |
| Core feature (play mode) | did not work | worked, with rough physics edges |
Read it honestly: this is one prompt and one comparison, and the harness run cost roughly 22 times more. The useful finding is not "harnesses are cheaper". It is that the model's ceiling was not the limit. The harness changed what the same weights could finish, and the remaining gaps (an unclear workflow, a wall the player could not jump past) were product judgment the harness had not been asked to check.
The same post holds the most transferable lesson. Separating the agent doing the work from the agent judging it was a strong lever, because tuning a standalone evaluator to be skeptical proved far more tractable than making a generator critical of its own work. The evaluator was given a browser tool so it could click through the live app before scoring, rather than grading a screenshot.
The five parts of a harness
The Learn Harness Engineering course organises the work into five subsystems. The diagram shows them as a ring around the model.
- Instructions tell the agent what to do and what to read first. Not one giant file: a short entry point that maps to deeper documents the agent opens when it needs them. OpenAI's harness-engineering write-up makes the same argument, treating AGENTS.md as a table of contents for a structured docs folder rather than an encyclopedia.
- Tools are the agent's hands: MCP servers, a shell, a web fetch, a database client. Each needs a scope and a clear failure signal, because a tool that fails without an error sends the agent reasoning over nothing.
- State survives the context window: a progress log, a feature list and the git history, all on disk, so the next session starts where the last one stopped.
- Verification is the only evidence that counts: tests, lint, type checks, a smoke run, or a separate evaluator agent. The agent may not declare victory without it.
- Scope and session lifecycle hold the agent to one feature at a time, start each session with an init script, and end it in a state someone could merge.
One gotcha before you adopt the vocabulary: the course itself uses two lists. The main framework names instructions, state, verification, scope and session lifecycle, while its frontier-product breakdowns (Claude Code, Codex, Pi and DeepSeek) use instructions, tools, environment, state and feedback. They describe the same ground at different cuts. Pick one list for your team and stick to it.
What a harness looks like on disk
Anthropic's earlier post on harnesses for long-running agents (26 November 2025) gives the concrete version. The first session runs an initializer agent with a different prompt. It writes an init.sh, a progress file and an initial git commit, and expands the user's request into a feature list. In their example that list ran past 200 features, every one marked failing. Every later session runs a coding agent that reads the progress file and git log, picks one failing feature, implements it, verifies it, commits with a descriptive message and updates the log.
repo/
AGENTS.md # about a page: what this is, where to look, how done is decided
docs/ # architecture, guardrails, decisions; opened on demand
feature_list.json # every feature, each "passing": false until verified
progress.md # what the last session did and what is next
init.sh # install, start services, run the smoke test
Two failure modes motivated that shape, and you will see both in customer repositories. Early on, the agent tries to build everything in one pass and runs out of context half way through, leaving a broken tree. Later, a fresh agent looks around, sees a lot of code and declares the project finished. The one-feature rule fixes the first; the all-failing feature list fixes the second, because "done" becomes a property of the list, not the agent's mood.
The repository that runs this site works this way. Its AGENTS.md is a short page of binding rules that points to separate documents for guardrails, architecture, design decisions and the content standard, and the build gates are the definition of done. Agents that skip a gate get caught by the gate, not by a reviewer reading their summary.
Learn Harness Engineering: the course
The walkinglabs repository is a project-based course: 14 lectures, 8 projects, translations into 15 languages, MIT licensed. Its strength is that each module builds a working artifact rather than summarising ideas. A harness-creator skill scaffolds the files above for your own project. Lecture 13 covers loop engineering (a maker agent and a checker agent run on a goal or a timer), and Lecture 14 covers graph engineering, for when one loop turns into several nodes with shared state and routing. The August 2026 breakdowns apply the five-part frame to how four frontier products build their harnesses.
Read it as a structured synthesis of public material. Its core references are the OpenAI and Anthropic posts, and its headline numbers come from them. When the README says Anthropic ran a "controlled experiment", go to the source: it was one prompt run two ways, which is useful evidence but not a controlled study.
MemoHarness: a harness that learns from its own runs
MemoHarness (arXiv 2607.14159, July 2026, ten authors) asks a sharper research question. Most teams ship one fixed harness for every task. Could the harness instead be tuned from execution history, and specialised per task, without labels at test time?
The paper splits a harness into six editable dimensions:
| Dimension | What it controls | Example edits |
|---|---|---|
| D1 Context assembly | What goes into the model input | restructure the prompt, add examples, compress context |
| D2 Tool interaction | When and how tools and retrieval run | enable retrieval, set top-k, rerank evidence |
| D3 Generation control | Decoding settings | raise the token budget, lower temperature, sample candidates |
| D4 Orchestration | The sequence of model calls | one call becomes plan, execute, refine |
| D5 Memory management | What persists across calls | keep state, summarise the trace, drop stale context |
| D6 Output processing | Turning raw output into an answer | extract, validate a schema, choose a fallback |
In phase A, a controller searches over edits to those dimensions on labelled cases. It ranks by task reward first, with token cost only as a tiebreaker, and stores per-case diagnoses plus distilled global patterns (what works, what fails, how they interact) in a dual-layer experience bank. In phase B, a new unlabelled case retrieves similar cases and patterns from the frozen bank, and the learned global harness is adapted to that case before one execution, with no test-time labels, feedback or extra search rounds.
The README's result figure is the honest part worth reading. On FinanceAgent, success rose from 42.5 percent to a 65.0 percent peak around search iterations 8 and 9. On LiveCodeBench it saturated almost immediately near the base model's ceiling and then oscillated within about a 4-point band. So harness search helped most where the task needed multi-step tool use and judgment, and barely mattered where the model alone was already near its limit. The authors state plainly that broader claims about statistical robustness and which component caused the gain are left to future work.
Practical notes before you clone it. It runs agents in a Harbor plus Daytona sandbox environment and needs a model provider key. And as of 7 October 2026 the repository has no license file. Without a license you can read the code and learn from the method, but you have no granted right to copy it into a customer's codebase.
Gotchas from running harnesses in practice
- The evaluator is generous by default. A model grading model output leans kind. Calibrate the evaluator with a few scored examples, give it a way to run the artifact (a browser, a test runner), and keep the final gate deterministic wherever you can.
- Agents edit the scoreboard. METR reported in June 2025 that frontier models, given coding tasks, tried to raise their scores by modifying the tests or the scoring code. Protect the feature list and the checks: the agent updates a "passing" flag only through the verification command, and changes to test files get human review.
- Silent tool failures look like reasoning. A web fetch that returns a bot-wall page, or an MCP server that returns an empty list on a timeout, gives the model plausible nothing to work with. Make tools return explicit errors, and log tool results next to model output so you can see which one failed.
- Context runs out before the task does. Long tasks need either compaction or clean resets with state on disk. Anthropic's harnesses used resets with an earlier model and could drop them with a later one, which is a reason to keep state on disk either way: you can change the model without changing the harness.
- Cost moves with reliability. The 9 versus 200 dollar comparison is the shape to expect. Set a budget per task and per session, and measure cost per verified feature, not cost per run.
- Instructions rot. An AGENTS.md written in week one describes a repository that no longer exists in week six. Treat it like code: short, reviewed, and updated in the same change that makes it wrong.
The FDE lens: setting up a harness at a customer
Forward deployed engineers meet this as a concrete ask: "we bought coding agent seats, why is the output unreliable on our monorepo?" The answer is rarely a different model. A useful first week looks like this:
- Day 1: read the repository the way an agent would. Can a stranger find the build command, the test command and the definition of done in under five minutes? If not, write the short AGENTS.md that says so and points to the docs that exist.
- Day 2: make verification one command. Most enterprise repos have tests that need a database, a VPN or a secret. A harness without a runnable check is a suggestion box. Containerise the test run or provide a fixture path.
- Day 3: add state. A progress file and a feature list for the current initiative, committed with the code, so sessions and people can hand off.
- Day 4: scope the tools. Which MCP servers, which credentials, which directories are writable. Agent sandboxing and least privilege are the customer's security team's first questions, so answer them before they are asked.
- Day 5: measure. Ten real tasks from the backlog, run with and without the harness, scored by the tests, with cost recorded. That table is what earns the next phase.
This is also an interview topic now. "How would you make a coding agent reliable inside a customer's legacy codebase?" is an agent system design question, and the strong answer names verification, durable state, scoped tools and a measured baseline before it names any model. The agentic evals and trajectories concept covers how to score the runs, context window management for FDE agents covers the state problem, and the agent vs workflow concept covers when you should not give the model the loop at all. For the context half of the story, see why context engineering replaced prompt engineering, and practise the design questions in RAG and agents and system design.

Turn it into offers. Work the real questions and concepts this maps to:
FAQ
Harness engineering is the design of everything around a language model that turns it into an agent you can rely on: the instructions it reads, the tools it can call, the state it keeps between sessions, the checks that decide whether work is done, and the limits on what it may touch. The model writes the code; the harness decides when, where and how, and whether the result counts.
Discussion (5)
The single highest-value line in all of this for customer work is that the evaluator should be a separate agent tuned to be skeptical. I have watched agents grade their own output as excellent for an entire afternoon. Splitting the judge out, and giving it the ability to actually run the thing, is the cheapest reliability win I know.
Agreed, with one caution from the same Anthropic post: the evaluator is still a model and still leans generous. Calibrate it with a handful of scored examples, and keep the final gate a deterministic check (tests, a schema, a row in a database) wherever one exists.
Reader question: if I can only build one piece of a harness this week, which one?
Verification. A command the agent must run, that a human trusts, and that fails loudly. Instructions and state make the agent efficient; verification is what stops it from telling you something works when it does not. Everything else can come next week.
The feature list starting as all failing is underrated. It flips the default from the agent deciding when it is finished to the list deciding, and it gives the customer a progress view they can read without asking anyone.
