FDEInterviews logo

You have two days, no context, and everybody has a different theory

Build a bounded rescue assessment from a supplied freight-deployment evidence packet. Separate containment from diagnosis, test competing explanations, distinguish stakeholder concerns from technical findings, and make the next commitment depend on evidence and access.

30 MIN

TL;DR: Contain active harm when needed, then build an assessment that separates observed facts, tested mechanisms and unresolved explanations. Agree a near-term update you can deliver, and make repair forecasts conditional on evidence, access and the customer's change process.

Where you are. You are joining an engagement that already has software, history and disappointed expectations. The task is to establish its current state and a feasible next decision, not to prove that the previous team was wrong.

Diagnose the workflow and the working relationship

In this fictional freight engagement, a five-month-old deployment has not reached its intended operating scope. Operators still resolve many exceptions manually, the original engineer now supports another account, and the sponsor wants a decision about continuing. Those facts describe a situation to investigate; they do not establish the cause or whether the software is salvageable.

Technical performance, workflow fit and confidence in delivery can all need attention. They influence one another, but each needs evidence. A working parser does not establish that users have access, the workflow is useful or the customer accepts the remaining support burden.

DimensionEvidence to seekAvoid inferring
System behaviorVersioned inputs/outputs, failures, reproduction and current configurationThe loudest complaint must name the mechanism
Workflow fitWork attempted, assistance, unresolved cases and operator constraintsA successful demo means routine work is covered
Delivery confidenceMissed/kept commitments, unanswered decisions and stated concernsA missed meeting or delegated attendee proves loss of trust
FeasibilityData, access, owners, change windows and remaining effortThe problem must be small because a new engineer found it

Ask the sponsor what decision they need and how progress will be judged. Ask operators to show examples of unfinished work. Ask the technical owners which access and recovery controls are available. A delegate may be the right decision-maker; establish their authority instead of diagnosing the relationship from seniority or attendance.

Containment can precede a complete diagnosis

If the system is sending incorrect notifications, corrupting data or exposing private information, waiting two days to finish an assessment can increase harm. Use the authorized incident process to reduce exposure: pause the affected action, route work to a staffed fallback or apply a verified reversible mitigation. Preserve useful evidence when feasible without making evidence collection a prerequisite for urgent containment.

Google's incident-response chapter describes stopping impact before completing root-cause investigation. It also shows why mitigation and understanding the underlying fault are separate tasks. Apply the principle to the actual risk and available controls; a rollback can itself be unsafe when state has changed.

A well-understood fix can be appropriate on arrival. Record the trigger, evidence, scope, validation, owner and recovery plan. Conversely, an impressive-looking patch with no reproduction or acceptance test is not a substitute for diagnosis. The clock alone decides neither case.

Use the first two days as an assessment cadence

DAY 1 Morning · listen, separately sponsor, operators, their IT, your team Afternoon · verify evidence logs, data, deploy history, the code DAY 2 Morning · hypothesis, then test it one query or one run that could refute it Afternoon · write and deliver assessment with confidence levels At any point: contain active harm and verify the result. A supported, authorized repair can happen early. Keep its evidence and recovery plan. The assessment records what is known, uncertain, changed and required next.

The two-day outline is a planning example. Access restrictions, incident urgency or source retention can change the order. If a required export will arrive later, deliver a partial assessment on the agreed date and say what cannot yet be established. Do not promise an exact diagnosis date that depends on unavailable evidence.

Speak with the sponsor, operators, customer technical owners and your own delivery team. Separate conversations can make it easier to hear constraints, while a joint review can resolve contradictory definitions. Neither format guarantees candor. Ask for the workflow, expected behavior, a recent concrete failure and changes around its onset.

Keep a small hypothesis register. Choose the next test by how much it can distinguish plausible explanations at acceptable risk and cost. One active experiment can keep the work focused; a rule allowing only one hypothesis encourages tunnel vision.

Grade the supplied evidence packet

The following packet shares the D1–D5 freight example used in this module's diagnosis lab. D1–D5 are complete synthetic UTC days, not dates from a real customer incident.

EvidenceSupplied statementWhat it supports
E1D1: 1,000 distinct inputs, 20 failed; D2: 2,000 inputs, 40 failedAbsolute failures doubled while both rates remain 2%
E2D3: 1,000 inputs, 220 failed; source format S2 begins at D3A 22% failure rate coincides with the recorded format change
E3Application B begins at D4; D3 still used AB's recorded deployment is later than this observed rate increase
E4Fixed application A accepts a string region and rejects the same code nested in an objectA reproduced parser-shape mechanism on a paired record
E5D3–D5 breakdown: 500 nested recognized codes, 25 genuinely unknown codes, 25 bad datesA conditional 500/550 recovery target for a shape repair
E6No destination-side action ledger has been providedReplay safety for external effects remains unestablished

E2 supplies a correlation: two things changed on the same day, and the packet cannot say which caused which. E4 demonstrates one mechanism, because it holds everything else constant and changes only the shape of one input, which is what separates a discriminating test from a coincidence in a timeline. E5 predicts a bounded recovery opportunity only if the full selected population matches that breakdown and passes replay. None establishes safety to resend every failed-looking record into production.

Practice with three statements:

  1. “Application B caused the onset.” The supplied timeline contradicts that narrow onset claim, assuming the deployment records are complete and correctly aligned. B could still introduce another problem later.
  2. “The compatible parser will fix 95.5% of failures.” Unsupported. Unknown-region labels cover 525/550 failures, which is 95.5 percent, but only 500 are the supplied shape-compatible subset: 500 ÷ 550 is about 90.9% of failures.
  3. “The model needs replacement.” Not established by this packet. Test the parser path and retain any independent model-quality concern as a separate hypothesis.

A confidence label should describe its basis. “Reproduced on one paired input; population recovery pending full replay” is more useful than an unexplained “90% confident.” It also tells the next engineer what evidence would strengthen or weaken the claim.

Choose tests that can change the decision

HypothesisDiscriminating evidenceWhat would change the assessment
The S2 shape triggers parser rejectionSame logical input and fixed app, change only the region representationIf both shapes pass, investigate the production path/configuration
Load explains the failure increaseArrival bursts, service times and resource saturation on aligned windowsA repeatable load failure not explained by input shape adds another mechanism
Model quality caused unresolved casesPaired reviewed outputs after successful ingestionBad answers among correctly parsed cases justify a separate model investigation
Failed-looking work already caused external effectsDestination receipts keyed to the selected records/actionsApplied or unknown effects change the replay plan

Do the safe local replay before changing a production parser. Ask for the highest-value missing evidence with a precise scope and purpose. A request for “all logs from the last six months” can be slow, expensive and unnecessarily sensitive; start with the failing interval and representative identifiers, then expand as justified.

Write an assessment that travels accurately

reportSynthetic D1–D5 packet | assessment example, not a customer incident

Rescue assessment: parser mechanism confirmed, full recovery still conditional

Current impact

D3 failure rate is 22% of distinct received records, compared with 2% on D1 and D2. The supplied D3–D5 set contains 550 failed records out of 2,500 inputs.

Established mechanism

Application A rejects the new nested region representation in a paired local replay. The recorded application B deployment occurs later than the observed onset. [1]

Proposed next step

Confirm the source contract and replay all 550 selected failures in an isolated output store with external actions denied. The supplied breakdown predicts 500 shape-compatible recoveries and 50 residual exceptions. Record the result rather than promising it in advance.

Unresolved boundary

We do not yet have destination action receipts, peak-window behavior or evidence for other source versions. Production replay waits for effect reconciliation and the approved change process. [2]

Next commitment

Deliver the isolated replay report at the agreed checkpoint, conditional on receipt of the selected export. If access is delayed, deliver the completed checks and the missing-evidence list at that checkpoint.

  1. [1] This distinguishes a reproduced mechanism from a broader causal claim about every failure.
  2. [2] An unknown external effect is a recovery constraint, not permission to try the action again.
Illustrative artifact from a fictional composite engagement, written as teaching material.

Report containment and repairs already performed, if any. Keep the business consequence visible: affected work, manual fallback capacity and outstanding decisions. A technically precise memo that never says what operators can do next leaves the rescue incomplete.

Repair commitments without inventing a relationship diagnosis

Agree the update cadence and audience. Daily notes can help during active recovery, but send useful changes, decisions and blockers rather than making a universal fourteen-day email rule. Record what you promised, what happened and what changed the forecast.

Discuss past defects with evidence and respect. Blameless reporting does not mean concealing a known error or refusing to explain how it affected the customer. Describe contributing conditions, impacts and corrective work without unsupported judgments about motives or competence. Google's postmortem guidance gives a concrete example of that approach.

A small visible improvement is worthwhile when it helps the workflow and fits the risk/effort budget. It cannot justify deferring a serious integrity problem or pretending that a cosmetic fix establishes recovery. Choose something the operator can verify, such as a reconciled unresolved-work queue with accurate reasons.

Recommend stopping when the evidence supports it

The use case may lack value, required access may be unavailable, or a safe service may cost more than the agreed scope can support. State the constraint, alternatives examined and remaining uncertainty. A narrower option is useful if supported by evidence; do not invent a viable replacement just to make a stop recommendation sound positive.

Stopping can itself require data export, retained evidence, outstanding-effect reconciliation and an agreed transition to manual work. Name those obligations. The decision belongs with the authorized product, customer and commercial owners; your assessment should make it reviewable.

Contain first, then build the assessment 1 Arrive with no context and three competing theories 2 Active harm? Contain before the diagnosis is done 3 Ask for the decision the sponsor actually needs 4 Grade the evidence what each item supports 5 Keep a hypothesis register small, and written down 6 Run the discriminating test safe replay before production 7 Observed, inferred, unknown in three labeled columns 8 Forecast, conditionally on evidence and on access Pause the affected action, route work to a staffed fallback, or apply a verified reversible mitigation. Preserve evidence where feasible, but never make evidence collection a prerequisite for urgent containment. A rollback can itself be unsafe once state has changed. Choose the next test by how much it can distinguish plausible explanations at acceptable risk and cost. One active experiment keeps the work focused; a rule allowing only one hypothesis encourages tunnel vision. "Reproduced on one paired input, population recovery pending full replay" is more useful than an unexplained "90 percent confident", because it says what would change the claim. If a required export arrives later, deliver a partial assessment on the agreed date and say what cannot yet be established. Do not promise a diagnosis date that depends on evidence nobody has.

The spine above is the two days as a sequence, with containment sitting before diagnosis rather than after it.

Do this before moving on

Write a one-page assessment from E1–E6 before reading the example memo again. Label observations, inferences, competing explanations and unknowns. Reproduce the 2%, 22%, 95.5% and 90.9% figures with their denominators, and explain which one is relevant to the parser recovery target.

Now add E7: an authorized operator confirms that duplicate notifications are still being emitted. Add immediate containment and effect reconciliation to the plan; do not wait for the end-of-day assessment. State what would prove the notifications have stopped and what might still be queued remotely.

Finally, change D2 to 440 failures out of 2,000 inputs. The 22% rate now precedes the recorded S2 change. Revise the onset claim and request earlier inputs/change evidence. The paired parser defect remains demonstrated, but it no longer explains the whole timing story by itself.

Go deeper

Key takeaways

  • Contain active harm when needed; a two-day assessment cadence is not a ban on early verified repairs.
  • Separate system behavior, workflow fit and stated delivery concerns instead of inferring trust from attendance.
  • Keep competing explanations and use tests that can change the next decision.
  • Distinguish reproduced mechanisms from population recovery predictions and unknown external effects.
  • Make commitments reviewable, conditional where necessary and explicit about the next owner or decision.

Check yourself

Answer before you look. Recalling it is what makes it stick; recognising it does not.

  1. 1Duplicate notifications are still being sent while you prepare the assessment. What takes priority?

  2. 2Why is 500/550 the supplied parser recovery target?

  3. 3A sponsor sends a delegate. What can you conclude?

  4. 4D2 is revised to a 22% failure rate before the recorded S2 rollout. What changes?

Sign in to track which lessons you have finished.