You have two days, no context, and everybody has a different theory
Build a bounded rescue assessment from a supplied freight-deployment evidence packet. Separate containment from diagnosis, test competing explanations, distinguish stakeholder concerns from technical findings, and make the next commitment depend on evidence and access.
30 MIN
TL;DR: Contain active harm when needed, then build an assessment that separates observed facts, tested mechanisms and unresolved explanations. Agree a near-term update you can deliver, and make repair forecasts conditional on evidence, access and the customer's change process.
Where you are. You are joining an engagement that already has software, history and disappointed expectations. The task is to establish its current state and a feasible next decision, not to prove that the previous team was wrong.
Diagnose the workflow and the working relationship
In this fictional freight engagement, a five-month-old deployment has not reached its intended operating scope. Operators still resolve many exceptions manually, the original engineer now supports another account, and the sponsor wants a decision about continuing. Those facts describe a situation to investigate; they do not establish the cause or whether the software is salvageable.
Technical performance, workflow fit and confidence in delivery can all need attention. They influence one another, but each needs evidence. A working parser does not establish that users have access, the workflow is useful or the customer accepts the remaining support burden.
| Dimension | Evidence to seek | Avoid inferring |
|---|---|---|
| System behavior | Versioned inputs/outputs, failures, reproduction and current configuration | The loudest complaint must name the mechanism |
| Workflow fit | Work attempted, assistance, unresolved cases and operator constraints | A successful demo means routine work is covered |
| Delivery confidence | Missed/kept commitments, unanswered decisions and stated concerns | A missed meeting or delegated attendee proves loss of trust |
| Feasibility | Data, access, owners, change windows and remaining effort | The problem must be small because a new engineer found it |
Ask the sponsor what decision they need and how progress will be judged. Ask operators to show examples of unfinished work. Ask the technical owners which access and recovery controls are available. A delegate may be the right decision-maker; establish their authority instead of diagnosing the relationship from seniority or attendance.
Containment can precede a complete diagnosis
If the system is sending incorrect notifications, corrupting data or exposing private information, waiting two days to finish an assessment can increase harm. Use the authorized incident process to reduce exposure: pause the affected action, route work to a staffed fallback or apply a verified reversible mitigation. Preserve useful evidence when feasible without making evidence collection a prerequisite for urgent containment.
Google's incident-response chapter describes stopping impact before completing root-cause investigation. It also shows why mitigation and understanding the underlying fault are separate tasks. Apply the principle to the actual risk and available controls; a rollback can itself be unsafe when state has changed.
A well-understood fix can be appropriate on arrival. Record the trigger, evidence, scope, validation, owner and recovery plan. Conversely, an impressive-looking patch with no reproduction or acceptance test is not a substitute for diagnosis. The clock alone decides neither case.
Use the first two days as an assessment cadence
The two-day outline is a planning example. Access restrictions, incident urgency or source retention can change the order. If a required export will arrive later, deliver a partial assessment on the agreed date and say what cannot yet be established. Do not promise an exact diagnosis date that depends on unavailable evidence.
Speak with the sponsor, operators, customer technical owners and your own delivery team. Separate conversations can make it easier to hear constraints, while a joint review can resolve contradictory definitions. Neither format guarantees candor. Ask for the workflow, expected behavior, a recent concrete failure and changes around its onset.
Keep a small hypothesis register. Choose the next test by how much it can distinguish plausible explanations at acceptable risk and cost. One active experiment can keep the work focused; a rule allowing only one hypothesis encourages tunnel vision.
Grade the supplied evidence packet
The following packet shares the D1–D5 freight example used in this module's diagnosis lab. D1–D5 are complete synthetic UTC days, not dates from a real customer incident.
| Evidence | Supplied statement | What it supports |
|---|---|---|
| E1 | D1: 1,000 distinct inputs, 20 failed; D2: 2,000 inputs, 40 failed | Absolute failures doubled while both rates remain 2% |
| E2 | D3: 1,000 inputs, 220 failed; source format S2 begins at D3 | A 22% failure rate coincides with the recorded format change |
| E3 | Application B begins at D4; D3 still used A | B's recorded deployment is later than this observed rate increase |
| E4 | Fixed application A accepts a string region and rejects the same code nested in an object | A reproduced parser-shape mechanism on a paired record |
| E5 | D3–D5 breakdown: 500 nested recognized codes, 25 genuinely unknown codes, 25 bad dates | A conditional 500/550 recovery target for a shape repair |
| E6 | No destination-side action ledger has been provided | Replay safety for external effects remains unestablished |
E2 supplies a correlation: two things changed on the same day, and the packet cannot say which caused which. E4 demonstrates one mechanism, because it holds everything else constant and changes only the shape of one input, which is what separates a discriminating test from a coincidence in a timeline. E5 predicts a bounded recovery opportunity only if the full selected population matches that breakdown and passes replay. None establishes safety to resend every failed-looking record into production.
Practice with three statements:
- “Application B caused the onset.” The supplied timeline contradicts that narrow onset claim, assuming the deployment records are complete and correctly aligned. B could still introduce another problem later.
- “The compatible parser will fix 95.5% of failures.” Unsupported. Unknown-region labels cover 525/550 failures, which is 95.5 percent, but only 500 are the supplied shape-compatible subset: 500 ÷ 550 is about 90.9% of failures.
- “The model needs replacement.” Not established by this packet. Test the parser path and retain any independent model-quality concern as a separate hypothesis.
A confidence label should describe its basis. “Reproduced on one paired input; population recovery pending full replay” is more useful than an unexplained “90% confident.” It also tells the next engineer what evidence would strengthen or weaken the claim.
Choose tests that can change the decision
| Hypothesis | Discriminating evidence | What would change the assessment |
|---|---|---|
| The S2 shape triggers parser rejection | Same logical input and fixed app, change only the region representation | If both shapes pass, investigate the production path/configuration |
| Load explains the failure increase | Arrival bursts, service times and resource saturation on aligned windows | A repeatable load failure not explained by input shape adds another mechanism |
| Model quality caused unresolved cases | Paired reviewed outputs after successful ingestion | Bad answers among correctly parsed cases justify a separate model investigation |
| Failed-looking work already caused external effects | Destination receipts keyed to the selected records/actions | Applied or unknown effects change the replay plan |
Do the safe local replay before changing a production parser. Ask for the highest-value missing evidence with a precise scope and purpose. A request for “all logs from the last six months” can be slow, expensive and unnecessarily sensitive; start with the failing interval and representative identifiers, then expand as justified.
Write an assessment that travels accurately
Rescue assessment: parser mechanism confirmed, full recovery still conditional
Current impact
D3 failure rate is 22% of distinct received records, compared with 2% on D1 and D2. The supplied D3–D5 set contains 550 failed records out of 2,500 inputs.
Established mechanism
Application A rejects the new nested region representation in a paired local replay. The recorded application B deployment occurs later than the observed onset. [1]
Proposed next step
Confirm the source contract and replay all 550 selected failures in an isolated output store with external actions denied. The supplied breakdown predicts 500 shape-compatible recoveries and 50 residual exceptions. Record the result rather than promising it in advance.
Unresolved boundary
We do not yet have destination action receipts, peak-window behavior or evidence for other source versions. Production replay waits for effect reconciliation and the approved change process. [2]
Next commitment
Deliver the isolated replay report at the agreed checkpoint, conditional on receipt of the selected export. If access is delayed, deliver the completed checks and the missing-evidence list at that checkpoint.
- [1] This distinguishes a reproduced mechanism from a broader causal claim about every failure.
- [2] An unknown external effect is a recovery constraint, not permission to try the action again.
Report containment and repairs already performed, if any. Keep the business consequence visible: affected work, manual fallback capacity and outstanding decisions. A technically precise memo that never says what operators can do next leaves the rescue incomplete.
Repair commitments without inventing a relationship diagnosis
Agree the update cadence and audience. Daily notes can help during active recovery, but send useful changes, decisions and blockers rather than making a universal fourteen-day email rule. Record what you promised, what happened and what changed the forecast.
Discuss past defects with evidence and respect. Blameless reporting does not mean concealing a known error or refusing to explain how it affected the customer. Describe contributing conditions, impacts and corrective work without unsupported judgments about motives or competence. Google's postmortem guidance gives a concrete example of that approach.
A small visible improvement is worthwhile when it helps the workflow and fits the risk/effort budget. It cannot justify deferring a serious integrity problem or pretending that a cosmetic fix establishes recovery. Choose something the operator can verify, such as a reconciled unresolved-work queue with accurate reasons.
Recommend stopping when the evidence supports it
The use case may lack value, required access may be unavailable, or a safe service may cost more than the agreed scope can support. State the constraint, alternatives examined and remaining uncertainty. A narrower option is useful if supported by evidence; do not invent a viable replacement just to make a stop recommendation sound positive.
Stopping can itself require data export, retained evidence, outstanding-effect reconciliation and an agreed transition to manual work. Name those obligations. The decision belongs with the authorized product, customer and commercial owners; your assessment should make it reviewable.
The spine above is the two days as a sequence, with containment sitting before diagnosis rather than after it.
Do this before moving on
Write a one-page assessment from E1–E6 before reading the example memo again. Label observations, inferences, competing explanations and unknowns. Reproduce the 2%, 22%, 95.5% and 90.9% figures with their denominators, and explain which one is relevant to the parser recovery target.
Now add E7: an authorized operator confirms that duplicate notifications are still being emitted. Add immediate containment and effect reconciliation to the plan; do not wait for the end-of-day assessment. State what would prove the notifications have stopped and what might still be queued remotely.
Finally, change D2 to 440 failures out of 2,000 inputs. The 22% rate now precedes the recorded S2 change. Revise the onset claim and request earlier inputs/change evidence. The paired parser defect remains demonstrated, but it no longer explains the whole timing story by itself.
Go deeper
- Diagnosing intermittent timeouts without customer code practices narrow evidence requests across a restricted access boundary.
- Proving what happened builds the record-level timeline and denominator checks behind this assessment.
- Salvaging a pilot that missed its metric connects technical findings to a bounded continuation decision.
- An executive escalation when scope and timeline broke rehearses communicating the impact and required decision without blame.
- Stakeholder management identifies the owners and authority needed to act on findings.
- Orienting in someone else's system supplies a technical exploration method for unfamiliar deployments.
Key takeaways
- Contain active harm when needed; a two-day assessment cadence is not a ban on early verified repairs.
- Separate system behavior, workflow fit and stated delivery concerns instead of inferring trust from attendance.
- Keep competing explanations and use tests that can change the next decision.
- Distinguish reproduced mechanisms from population recovery predictions and unknown external effects.
- Make commitments reviewable, conditional where necessary and explicit about the next owner or decision.
Check yourself
Answer before you look. Recalling it is what makes it stick; recognising it does not.
1Duplicate notifications are still being sent while you prepare the assessment. What takes priority?
2Why is 500/550 the supplied parser recovery target?
3A sponsor sends a delegate. What can you conclude?
4D2 is revised to a 22% failure rate before the recorded S2 rollout. What changes?
Sign in to track which lessons you have finished.
