Practice the job, on a box that answers back.
Every mission drops you on a customer's machine with a problem, logs, a contract and a terminal. You find the evidence, change the thing, and prove it. Objectives tick when the world changes, not when you type the expected command, and the debrief tells you what you can now say in the room. The same customers come back mission after mission, so what you learn compounds.
The brief tells you where to start; objectives arrive in order. The free missions are guided.
Objectives visible, method yours. Most premium missions; the interviewer-grade version of the guided one.
A clock, a broken system, and a diagnosis you commit to. The after-action shows the evidence path you took.
The simulator, no mission
free · no account · nothing scoredThe two customer systems the missions are set in, running normally. Inject a fault from the sandbox, watch it land in the topology, the traces, the metrics and the logs at once, and fix it from the terminal. When it clicks, the missions are the same systems with a problem that has already happened.
The Calder support agent system with no mission: gateway, runtime, router, providers, tools, queue and approvals. Inject faults, change config, let time pass.
Halvard's carrier integration with no mission: carrier webhooks, receiver, dedupe store, billing, pricing, rates API and the nightly sync.
Field basics
0/3 doneOrient on an unfamiliar box: logs, configs, the job that failed and the value that actually ran.
- FreeGuided · beginner · 20 minDay one on the boxHalvard Freight
Your first morning at Halvard Freight and the nightly rate sync failed. Orient on an unfamiliar machine, find the failure in the logs, work out which config value actually ran, and prove the fix with a dry run.
- Free accountChallenge · beginner · 20 minThe config that liesHalvard Freight
Halvard's receiver has three places a timeout can come from and they disagree. Work out which value actually runs, why the one in the repo is not it, and put the right value in the layer that wins.
- PremiumChallenge · intermediate · 25 minTwo clocksHalvard Freight
Halvard's nightly cut assigns the last hour of deliveries to the wrong day, but only some nights. One system stamps UTC, one Oslo local time, and the boundary moves twice a year. Find the seam, prove it with the data, and fix the cut.
API debugging
0/3 doneRetries, idempotency, rate limits, pagination and the 200 that failed: the integration bugs that page you at night.
- FreeGuided · beginner · 25 minThe webhook that fires twiceHalvard Freight
Halvard Freight's carrier webhook created two invoices for one delivery. Find the duplicate in the logs, explain why the carrier retried, and make the billing write safe to repeat.
- Free accountChallenge · intermediate · 25 min429 at nine o'clockMeridian Health
Every Monday at nine Meridian's claims assistant slows to a crawl and a third of runs fail. The provider is rate limiting, and the router's retries are making it worse. Read the storm, name the mechanism, and tame it with one config change.
- PremiumChallenge · intermediate · 30 minThe 200 that failedMeridian Health
Meridian's nightly claims sync reports success and the totals are wrong. The claims API returns errors inside 200 responses and a cursor the client ignores. Find both, fix the client's config, and make the dry run tell the truth.
Building agents
0/4 doneTurn budgets, tool schemas, structured output and nested retries: what makes a model loop safe to ship.
- PremiumGuided · intermediate · 35 minGive the agent a stop ruleCalder Equipment
Calder's support agent ran 47 turns on one case overnight, calling the same tool against a service in maintenance. Read the trace, bound the loop, handle the tool state the agent never handled, and replay the case to a clean handoff.
- PremiumChallenge · intermediate · 25 minStructured output, or nothingCalder Equipment
Calder's agent sometimes answers in prose where a proposal object is expected, and the pipeline counts those as successes with an empty action. Count them, make the output a schema or a hold, and stop success from meaning nothing happened.
- PremiumChallenge · advanced · 30 minThe model that asked twiceCalder Equipment
One slow provider call became nine at Calder because three layers each retried it. Read the trace, count the multiplication, give the deadline one owner, and prove a case now fails fast instead of nine times slowly.
- PremiumChallenge · intermediate · 30 minTool schemas that biteCalder Equipment
Calder's agent called read_support_case with a case id that does not exist, got a neighbour's record back, and proposed a replacement for it. Tighten the schema and the adapter, and replay the case to a clarification instead.
Durable actions
0/4 doneIntent before send, unknown outcomes, reconciliation, backlogs: the actions a retry must never repeat.
- PremiumGuided · advanced · 35 minKill the workerCalder Equipment
Calder's action worker sent a replacement request and died before recording the receipt. Reproduce the crash, find out what the service actually committed, choose a recovery that cannot ship a second pump, and prove it.
- PremiumChallenge · intermediate · 25 minPause writes, keep lookupsCalder Equipment
A bad release is proposing wrong replacements at Calder and it must stop shipping now, without stopping customers from getting answers. Pick the stop that blocks effects and not lookups, apply it, and prove both halves live.
- PremiumIncident · advanced · 30 minQueue backlog at 3amCalder Equipment
Calder's action queue has been filling for forty minutes and nothing is draining it. Contain it, find out why the consumers went away, get the backlog moving without shipping anything twice, and tell the morning shift what happened.
- PremiumChallenge · intermediate · 25 minThe retry that shipped two pumpsCalder Equipment
Calder shipped two pumps for one approved replacement after a network blip: the worker retried with a new key. Find the key shape that made the retry a new request, fix it, and show the service refusing a changed payload under the same key.
Evaluation and release
0/4 donePaired cases, judges, regression gates and drifting labels: the evidence that decides a release.
- FreeGuided · beginner · 20 minThe candidate that scores higherCalder Equipment
Calder's new agent release scores 9 of 11 against the old one's 8. Read the paired results case by case, find the case where the better score hides a worse action, and say what should happen to this release.
- PremiumChallenge · intermediate · 30 minCalibrate the judgeCalder Equipment
Calder's model judge scores the agent's proposals and it disagrees with the human reviewer on the cases that matter. Compare the two on the labelled set, name which way the judge is wrong, fix the rubric it was given, and rerun.
- PremiumChallenge · intermediate · 25 minGolden set driftMeridian Health
Meridian's evaluation passes every night and the assistant is wrong about the new coverage rules: the golden set was written under last year's policy. Find the labels that rotted, version the manifest, and make the run tell the truth.
- PremiumChallenge · intermediate · 30 minThe regression gateCalder Equipment
Calder's candidate release scores higher than the baseline and someone wants it shipped today. Read the paired packet, find the case that should block it, turn that failure into a release-blocking test, and make the gate say hold.
Serving and cost
0/4 doneModel routing, memory fit, cache keys and context budgets: the numbers behind a serving decision.
- PremiumChallenge · intermediate · 25 minDoes it fitNorthwind Insurance
Northwind wants the claims model on their own two 80 GB cards; the sheet says 70 GB of weights. Work out whether it fits with a KV cache at their concurrency, in the units the card uses, and set the admission limit that keeps it fitting.
- PremiumChallenge · intermediate · 25 minRoute the cheap trafficNorthwind Insurance
At Monday peak Northwind's primary model sits at its rate limit while the fallback idles. Find the bottleneck, route the traffic that does not need the large model, and keep the policy reasoning where it belongs.
- PremiumChallenge · intermediate · 25 minStale cache after a policy changeNorthwind Insurance
Northwind published a new coverage policy at nine and the assistant kept quoting the old one until lunch. Retrieval was fine. Find the layer that remembered, key it by what the answer depends on, and keep the hit rate.
- PremiumChallenge · intermediate · 25 minTruncated policyCalder Equipment
Calder's agent started proposing replacements without inspections after a policy update. Nothing in the prompt changed. The compiled context did: it is over budget and the clause that governs the decision is the one being cut.
Incidents
0/4 doneThe clock is running and the system is broken: contain, diagnose, commit, and defend it afterwards.
- PremiumIncident · advanced · 30 minIntermittent timeouts, no source accessNorthwind Insurance
Northwind's claims assistant is timing out for a third of customers and you cannot read their code. Contain it from the console, find the phase that is slow, commit to a root cause, and write the first status update.
- PremiumIncident · advanced · 35 minNothing alerts when quality dropsMeridian Health
Meridian's claims assistant has been quietly wrong for three days. Latency is fine, errors are zero, and the evaluation on an unchanged manifest fell from 9 to 6. Find what changed, contain it, and add the alert that should have fired.
- PremiumIncident · advanced · 45 minRelease review: defend the agentCalder Equipment
The capstone. Calder's support agent is up for release and the board has forty minutes. Gather evidence from the live system, answer the three questions a board asks, and commit to hold or advance with reasons a reviewer accepts.
- PremiumIncident · advanced · 40 minThe first 48 hoursHalvard Freight
Halvard's billing rejected every invoice since two in the morning and the carrier has been redelivering ever since. Contain the storm, find the rotated token, restore invoicing, and hand over evidence the next shift can act on.
What is behind the terminal
The box is a deterministic simulation, and it says so on the prompt. Files, logs and payloads are written for the mission; services answer from a contract; the programs you run (a worker, an agent, a receiver) behave according to the config and policy you edit, so a wrong fix produces a wrong result you can read. That is the point: the commands are real, the evidence is consistent, and a fix is proven by replaying the failure, the way you would on a real engagement. It is also why a thousand readers can be on it at once without a server doing anything for any of them.
Missions link into the concepts, questions and courses that explain the mechanism, before and after. Read first if you want; most people learn it faster by breaking it.
Questions
Is anything actually running?
A simulated box runs in your browser tab: a filesystem, environment, logs, services that answer HTTP and programs that respond to what you change. Nothing you type is sent to a server. The world is a recording, so every reader sees the same evidence and every fix can be checked exactly.
Which missions are free?
3 missions are free without an account, 3 more open with a free account, and the remaining 24, including every incident and challenge, are part of premium. A locked mission still shows its brief and objectives so you know what you are practising before you pay.
How is a mission checked?
Objectives are checked against the state of the box, never against the exact command you typed. If you solve it another way, it still counts. Hints come in three levels, from a question to a near-answer, and the debrief says what you can now claim in an interview.
Does progress save?
Sessions persist in your browser, so you can leave and come back. A signed-in reader can save a completion to their account; the server replays the transcript before it counts, so a saved completion means the mission was solved.
