FDEInterviews logo

Harness Engineering, Explained: What Makes AI Coding Agents Reliable (Learn Harness Engineering and MemoHarness)

Harness engineering is everything you build around a model so an agent finishes real work: instructions, tools, state, verification and scope. What Anthropic measured, what the Learn Harness Engineering course and the MemoHarness paper add, the gotchas, and how an FDE sets one up at a customer.

BY LUKAS HOFFMANN · FDEINTERVIEWS EDITORIAL · UPDATED OCTOBER 7, 2026 · 10 MIN READ

PRACTICE THIS:RAG and agent design questions ·Agentic evals and trajectories, the concept ·The Production AI Agents course ·The must-know FDE questions

Harness engineering is the work of building everything around a language model so that an agent finishes real tasks reliably: the instructions it reads, the tools it may call, the state it keeps on disk between sessions, the checks that decide whether work is done, and the limits on what it may touch. The model decides what code to write. The harness decides when, where and how, and whether the result counts. The term spread through late 2025 and 2026 as OpenAI and Anthropic published how they run coding agents for hours at a time, and two open-source projects now teach and test it: the Learn Harness Engineering course from walkinglabs (MIT, about 19,400 GitHub stars as of 7 October 2026) and the MemoHarness research code, which tries to learn harness settings from the agent's own runs. This post explains the pattern, what each project adds, the gotchas, and how a forward deployed engineer sets a harness up inside a customer's repository.

Same model, different harness: the evidence

The clearest data point comes from Anthropic's engineering post on harness design for long-running application development (24 March 2026). The author gave the same model the same one-sentence prompt, a 2D retro game editor, twice. Run alone, it worked for about 20 minutes, cost about 9 dollars, and produced an app whose play mode did not work. Run inside a three-agent harness (a planner, a generator and an evaluator), it worked for about 6 hours, cost about 200 dollars, and produced an editor where you could actually move a character and play a level.

Solo agentPlanner, generator and evaluator harness
Durationabout 20 minutesabout 6 hours
Costabout $9about $200
Scopewhat the model chose to builda planner expanded the prompt into a 16-feature spec over ten sprints
Core feature (play mode)did not workworked, with rough physics edges

Read it honestly: this is one prompt and one comparison, and the harness run cost roughly 22 times more. The useful finding is not "harnesses are cheaper". It is that the model's ceiling was not the limit. The harness changed what the same weights could finish, and the remaining gaps (an unclear workflow, a wall the player could not jump past) were product judgment the harness had not been asked to check.

The same post holds the most transferable lesson. Separating the agent doing the work from the agent judging it was a strong lever, because tuning a standalone evaluator to be skeptical proved far more tractable than making a generator critical of its own work. The evaluator was given a browser tool so it could click through the live app before scoring, rather than grading a screenshot.

The five parts of a harness

The Learn Harness Engineering course organises the work into five subsystems. The diagram shows them as a ring around the model.

Instructions short AGENTS.md as a map deeper docs read on demand definition of done Tools MCP servers, shell, fetch each with a scope and a limit failures surfaced, not hidden the model decides what to write State progress file on disk feature list, all start failing git history as the undo log Verification tests, lint, type checks separate skeptical evaluator stop only when it passes Scope and lifecycle one feature at a time init at start, clean exit The harness decides when, where and how, and whether the result counts.
  • Instructions tell the agent what to do and what to read first. Not one giant file: a short entry point that maps to deeper documents the agent opens when it needs them. OpenAI's harness-engineering write-up makes the same argument, treating AGENTS.md as a table of contents for a structured docs folder rather than an encyclopedia.
  • Tools are the agent's hands: MCP servers, a shell, a web fetch, a database client. Each needs a scope and a clear failure signal, because a tool that fails without an error sends the agent reasoning over nothing.
  • State survives the context window: a progress log, a feature list and the git history, all on disk, so the next session starts where the last one stopped.
  • Verification is the only evidence that counts: tests, lint, type checks, a smoke run, or a separate evaluator agent. The agent may not declare victory without it.
  • Scope and session lifecycle hold the agent to one feature at a time, start each session with an init script, and end it in a state someone could merge.

One gotcha before you adopt the vocabulary: the course itself uses two lists. The main framework names instructions, state, verification, scope and session lifecycle, while its frontier-product breakdowns (Claude Code, Codex, Pi and DeepSeek) use instructions, tools, environment, state and feedback. They describe the same ground at different cuts. Pick one list for your team and stick to it.

What a harness looks like on disk

Anthropic's earlier post on harnesses for long-running agents (26 November 2025) gives the concrete version. The first session runs an initializer agent with a different prompt. It writes an init.sh, a progress file and an initial git commit, and expands the user's request into a feature list. In their example that list ran past 200 features, every one marked failing. Every later session runs a coding agent that reads the progress file and git log, picks one failing feature, implements it, verifies it, commits with a descriptive message and updates the log.

repo/
  AGENTS.md            # about a page: what this is, where to look, how done is decided
  docs/                # architecture, guardrails, decisions; opened on demand
  feature_list.json    # every feature, each "passing": false until verified
  progress.md          # what the last session did and what is next
  init.sh              # install, start services, run the smoke test

Two failure modes motivated that shape, and you will see both in customer repositories. Early on, the agent tries to build everything in one pass and runs out of context half way through, leaving a broken tree. Later, a fresh agent looks around, sees a lot of code and declares the project finished. The one-feature rule fixes the first; the all-failing feature list fixes the second, because "done" becomes a property of the list, not the agent's mood.

The repository that runs this site works this way. Its AGENTS.md is a short page of binding rules that points to separate documents for guardrails, architecture, design decisions and the content standard, and the build gates are the definition of done. Agents that skip a gate get caught by the gate, not by a reviewer reading their summary.

Learn Harness Engineering: the course

The walkinglabs repository is a project-based course: 14 lectures, 8 projects, translations into 15 languages, MIT licensed. Its strength is that each module builds a working artifact rather than summarising ideas. A harness-creator skill scaffolds the files above for your own project. Lecture 13 covers loop engineering (a maker agent and a checker agent run on a goal or a timer), and Lecture 14 covers graph engineering, for when one loop turns into several nodes with shared state and routing. The August 2026 breakdowns apply the five-part frame to how four frontier products build their harnesses.

Read it as a structured synthesis of public material. Its core references are the OpenAI and Anthropic posts, and its headline numbers come from them. When the README says Anthropic ran a "controlled experiment", go to the source: it was one prompt run two ways, which is useful evidence but not a controlled study.

MemoHarness: a harness that learns from its own runs

MemoHarness (arXiv 2607.14159, July 2026, ten authors) asks a sharper research question. Most teams ship one fixed harness for every task. Could the harness instead be tuned from execution history, and specialised per task, without labels at test time?

The paper splits a harness into six editable dimensions:

DimensionWhat it controlsExample edits
D1 Context assemblyWhat goes into the model inputrestructure the prompt, add examples, compress context
D2 Tool interactionWhen and how tools and retrieval runenable retrieval, set top-k, rerank evidence
D3 Generation controlDecoding settingsraise the token budget, lower temperature, sample candidates
D4 OrchestrationThe sequence of model callsone call becomes plan, execute, refine
D5 Memory managementWhat persists across callskeep state, summarise the trace, drop stale context
D6 Output processingTurning raw output into an answerextract, validate a schema, choose a fallback
Phase A: search on labelled cases Labelled casestask + reference Candidate harnessedits over D1 to D6 Executeone run per case Scorereward, then cost Dual-layer experience bank per-case diagnoses + distilled global patterns Phase B: adapt per case, no labels New caseno reference answer Retrievecases and patterns Adapt harnessglobal to per-case Execute, answerone pass, no search

In phase A, a controller searches over edits to those dimensions on labelled cases. It ranks by task reward first, with token cost only as a tiebreaker, and stores per-case diagnoses plus distilled global patterns (what works, what fails, how they interact) in a dual-layer experience bank. In phase B, a new unlabelled case retrieves similar cases and patterns from the frozen bank, and the learned global harness is adapted to that case before one execution, with no test-time labels, feedback or extra search rounds.

The README's result figure is the honest part worth reading. On FinanceAgent, success rose from 42.5 percent to a 65.0 percent peak around search iterations 8 and 9. On LiveCodeBench it saturated almost immediately near the base model's ceiling and then oscillated within about a 4-point band. So harness search helped most where the task needed multi-step tool use and judgment, and barely mattered where the model alone was already near its limit. The authors state plainly that broader claims about statistical robustness and which component caused the gain are left to future work.

Practical notes before you clone it. It runs agents in a Harbor plus Daytona sandbox environment and needs a model provider key. And as of 7 October 2026 the repository has no license file. Without a license you can read the code and learn from the method, but you have no granted right to copy it into a customer's codebase.

Gotchas from running harnesses in practice

  1. The evaluator is generous by default. A model grading model output leans kind. Calibrate the evaluator with a few scored examples, give it a way to run the artifact (a browser, a test runner), and keep the final gate deterministic wherever you can.
  2. Agents edit the scoreboard. METR reported in June 2025 that frontier models, given coding tasks, tried to raise their scores by modifying the tests or the scoring code. Protect the feature list and the checks: the agent updates a "passing" flag only through the verification command, and changes to test files get human review.
  3. Silent tool failures look like reasoning. A web fetch that returns a bot-wall page, or an MCP server that returns an empty list on a timeout, gives the model plausible nothing to work with. Make tools return explicit errors, and log tool results next to model output so you can see which one failed.
  4. Context runs out before the task does. Long tasks need either compaction or clean resets with state on disk. Anthropic's harnesses used resets with an earlier model and could drop them with a later one, which is a reason to keep state on disk either way: you can change the model without changing the harness.
  5. Cost moves with reliability. The 9 versus 200 dollar comparison is the shape to expect. Set a budget per task and per session, and measure cost per verified feature, not cost per run.
  6. Instructions rot. An AGENTS.md written in week one describes a repository that no longer exists in week six. Treat it like code: short, reviewed, and updated in the same change that makes it wrong.

The FDE lens: setting up a harness at a customer

Forward deployed engineers meet this as a concrete ask: "we bought coding agent seats, why is the output unreliable on our monorepo?" The answer is rarely a different model. A useful first week looks like this:

  • Day 1: read the repository the way an agent would. Can a stranger find the build command, the test command and the definition of done in under five minutes? If not, write the short AGENTS.md that says so and points to the docs that exist.
  • Day 2: make verification one command. Most enterprise repos have tests that need a database, a VPN or a secret. A harness without a runnable check is a suggestion box. Containerise the test run or provide a fixture path.
  • Day 3: add state. A progress file and a feature list for the current initiative, committed with the code, so sessions and people can hand off.
  • Day 4: scope the tools. Which MCP servers, which credentials, which directories are writable. Agent sandboxing and least privilege are the customer's security team's first questions, so answer them before they are asked.
  • Day 5: measure. Ten real tasks from the backlog, run with and without the harness, scored by the tests, with cost recorded. That table is what earns the next phase.

This is also an interview topic now. "How would you make a coding agent reliable inside a customer's legacy codebase?" is an agent system design question, and the strong answer names verification, durable state, scoped tools and a measured baseline before it names any model. The agentic evals and trajectories concept covers how to score the runs, context window management for FDE agents covers the state problem, and the agent vs workflow concept covers when you should not give the model the loop at all. For the context half of the story, see why context engineering replaced prompt engineering, and practise the design questions in RAG and agents and system design.

THE ONE-PAGE VERSION
Infographic on harness engineering for AI coding agents: instructions as a short map with deeper docs on demand, state kept on disk in a feature list that starts failing and a progress file, verification by tests and a separate skeptical evaluator, and scope limited to one feature per session with an init step and a clean exit.
↧ DownloadShare on X ↗Share on LinkedIn ↗
PRACTICE THIS

Turn it into offers. Work the real questions and concepts this maps to:

FAQ

What is harness engineering?▲

Harness engineering is the design of everything around a language model that turns it into an agent you can rely on: the instructions it reads, the tools it can call, the state it keeps between sessions, the checks that decide whether work is done, and the limits on what it may touch. The model writes the code; the harness decides when, where and how, and whether the result counts.

How is harness engineering different from prompt engineering and context engineering?▼
Does a harness really change results, or just cost?▼
Is MemoHarness ready to use in production?▼
Where should I start learning harness engineering?▼

Discussion (5)

Arjun MehtaEditor

The single highest-value line in all of this for customer work is that the evaluator should be a separate agent tuned to be skeptical. I have watched agents grade their own output as excellent for an entire afternoon. Splitting the judge out, and giving it the ability to actually run the thing, is the cheapest reliability win I know.

Lukas HoffmannEditor

Agreed, with one caution from the same Anthropic post: the evaluator is still a model and still leans generous. Calibrate it with a handful of scored examples, and keep the final gate a deterministic check (tests, a schema, a row in a database) wherever one exists.

Cole SullivanContributor

Reader question: if I can only build one piece of a harness this week, which one?

Emily CarterEditor

Verification. A command the agent must run, that a human trusts, and that fails loudly. Instructions and state make the agent efficient; verification is what stops it from telling you something works when it does not. Everything else can come next week.

Brandon FosterEditor

The feature list starting as all failing is underrated. It flips the default from the agent deciding when it is finished to the list deciding, and it gives the customer a progress view they can read without asking anyone.