You have two days, no context, and everybody has a different theory
Being dropped into a failing engagement is a distinct skill from building one. Separate the technical failure from the trust failure, take evidence over accounts, and deliver a diagnosis with a confidence level rather than a fix: a confident wrong answer in week one ends the account faster than the original problem.
18 MIN
TL;DR: Spend the first day gathering evidence rather than opinions, establish what is actually true independently of what people say, and deliver a diagnosis that separates the technical problem from the trust problem. Do not promise a fix in the first 48 hours; promise an assessment, and hit that date exactly.
Where you are. Module six, and a change of situation. Every previous lesson assumed you were building. This one assumes somebody else built it, it is failing, and you have been sent in with a fortnight and a great deal of ambient anxiety.
Two failures, and only one is technical
The deployment is five months old. It was supposed to go live in eight weeks. The customer's sponsor has stopped attending the weekly call, the engineer who built it has moved to another account, and your own leadership has described the situation with a word like "salvage".
There are two problems in that paragraph and conflating them is the most common way a rescue fails.
The technical failure is whatever the system does or does not do. It is usually smaller than the reputation suggests, and it is usually not the thing everybody names.
The trust failure is that the customer no longer believes what they are told. This one is often larger, it decays further every week, and it is not repaired by fixing the software. A rescue that fixes the system and ignores the relationship produces working software the customer cancels anyway.
You need a read on both by the end of the first week, and they are diagnosed differently.
| The technical failure | The trust failure | |
|---|---|---|
| Diagnosed from | Logs, data, a reproduction | Who still takes the meeting |
| Usual size | Smaller than its reputation | Larger than anyone admits |
| Direction over time | Static once found | Decays every week you are quiet |
| Fixed by | Engineering | Prediction, then delivery, repeated |
| Signal it is real | You can reproduce it | The sponsor sends a delegate |
The last row on the right is the one to watch for in your first week. When a sponsor who used to attend personally starts sending someone junior, that is not a calendar clash. It is a forecast.
The first 48 hours, hour by hour
Listen first, separately, and to four groups. The sponsor who funded it, the operators who were meant to use it, their technical staff, and your own people. Separately matters: in a joint meeting everybody performs, and the operator will not say the thing they will say alone.
Ask each the same three questions. What was this supposed to do? What does it do now? When did you last believe it was going to work? That third question is the useful one, because the date people give is usually the same date, and whatever happened that week is where the story is.
Then verify independently. Every account you have just heard is sincere and partial. Before accepting any of it, establish from evidence: what is actually deployed, when it last changed, what it processed in the last month, and what fraction of that was correct. The next lesson is entirely about this, because it is where a rescue is won.
Form one hypothesis and try to refute it. Not five. One, stated plainly, with the specific query or test run that would prove it wrong. Rescues go badly when the new engineer arrives with a list of everything that is imperfect, which is indistinguishable from criticising the previous team and produces no decision.
Deliver an assessment, not a fix
At the end of day two you present something. What it must contain:
What is true. The state of the system, in facts with evidence attached. Not "the pipeline is unreliable" but "of 12,400 records last month, 3,100 were quarantined, 2,900 of those for one reason".
The diagnosis, with a confidence level. "Most likely X, and I am fairly confident; possibly also Y, which I have not been able to test yet." Stating confidence honestly is what makes you believable to a room that has been told certain things before and found them wrong.
What you have not established. The unknowns, named. This is counter-intuitive in a room that wants reassurance, and it is precisely what rebuilds credibility, because everybody there already knows the situation is uncertain and has been waiting for somebody to say so.
What happens next and when. A date for the plan, not the plan. Then hit that date exactly, because the first commitment you keep is worth more than the content of it.
What it must not contain is a promise about the fix. You do not yet know, and the room has already been promised things.
The judgment call nobody wants
Occasionally the honest diagnosis is that the deployment should not continue: the use case was wrong, the data cannot support it, or the value was never there.
Saying so is the most valuable thing an engineer can do in a rescue and the hardest, because it disappoints everybody in the room and appears to be your own company losing revenue. It is also, usually, cheaper for everyone than another two quarters of the same, and it frequently preserves the account for a different piece of work, because the customer remembers who told them the truth.
If you reach that conclusion, bring an alternative. "This should stop, and here is the narrower thing in the same account that would work" is a recommendation. "This should stop" alone is a resignation letter for the project.
What to do about the trust half
Running alongside the technical work from day one, and mostly consisting of small, boring reliability.
Say what you will do, then do exactly that, on the day you said. Send a short written update every single day of the first fortnight, even when the content is that you are still reading logs. Never explain the previous team's mistakes to the customer, which reads as self-serving and tells them nothing they can use. And find one small thing that visibly works within the first week, because five months of nothing improving has taught everyone that nothing improves.
That last one is not theatre. A single fixed irritation, shipped in week one, is the evidence that the pattern has changed, and it buys you the fortnight you need for the real work.
Do this before moving on
Take a project you have seen fail and reconstruct the 48-hour assessment you would have delivered on the day you arrived. Write the three sections: what is true with evidence, the diagnosis with a confidence level, and the named unknowns. Then check whether the thing everybody blamed at the time appears in your evidence section. It usually does not, and noticing that gap is the whole skill.
Go deeper
- Diagnosing intermittent timeouts without access to customer code is a rescue diagnosis in miniature.
- Salvaging a pilot that quietly missed its metric is the commercial half of this situation.
- An executive escalation when scope and timeline both broke is the conversation your assessment is preparing for.
- Stakeholder management is how you read which relationships survived the five months before you arrived.
- Orienting in someone else's system is the technical method, under time pressure here.
Key takeaways
- Separate the technical failure from the trust failure; fixing the first while ignoring the second produces software the customer cancels anyway.
- Interview four groups separately and ask when they last believed it would work, because that date locates the story.
- Verify everything independently: sincere accounts are partial, and the named cause is usually not the actual one.
- Deliver an assessment with confidence levels and explicit unknowns, not a fix, and hit the date you gave for the plan.
- Ship one visible improvement in week one; after five months of nothing changing, evidence that the pattern has changed is what buys you time.
Check yourself
Answer before you look. Recalling it is what makes it stick; recognising it does not.
1Why is delivering a fix within the first 48 hours of a rescue a mistake rather than an impressive start?
2You interview stakeholders and they broadly agree on what went wrong. Why still verify independently?
3Why include explicitly named unknowns in a day-two assessment, when the room wants reassurance?
Sign in to track which lessons you have finished.
