Leadership & Judgement

Good escalation is a technical skill

Why good escalation is a technical discipline built on evidence, timing, judgement, ownership and a clear request.

Leadership & Judgement escalationtroubleshootingincident responsetechnical communication

In view

  • Pillar: Leadership & Judgement
  • Maturity: carefully framed publication
  • Edited for publication and safe disclosure.

Editorial note

Carefully framed
  • Some examples are deliberately abstracted to keep the judgement useful without exposing private systems, people, weaknesses or operational detail.
  • Live incidents, internal escalation paths, supplier identities and support records.
  • Environment-specific logs, system details, timestamps or troubleshooting evidence.
  • Named staff, capability assessments or internal service-management weaknesses.

I used to think of escalation mainly as a support process.

You investigate something until you reach the limit of your access or knowledge, then you send it to somebody more senior or to another team.

I now think that description misses the important part.

The difficult part is not knowing that escalation exists.

The difficult part is recognising when to escalate, what evidence to carry with you and what you actually need from the person receiving it.

That makes escalation a technical skill.

There are two easy ways to get it wrong.

Escalate too quickly and another engineer spends time repeating checks that could reasonably have been completed already.

Escalate too late and an incident sits with somebody for hours while they experiment around a system they do not understand well enough to diagnose.

Neither is good ownership.

The useful point sits somewhere between them.

An escalation should move the investigation forwards

One of the simplest tests I use is this:

Does the next person know more after receiving the escalation than they would have known from the original fault report?

If the answer is no, the escalation probably needs more work.

“The network is down.”

“The application doesn’t work.”

“The supplier needs to investigate.”

“There is a firewall issue.”

Those statements may eventually prove true, but they do not tell another engineer much.

A useful escalation normally carries observations rather than assumptions.

What is affected?

What remains working?

When did the problem begin?

Can it be reproduced?

What changed recently?

What has already been checked?

Which test results are important?

What has been ruled out?

What is still uncertain?

That changes the quality of the next investigation.

The receiving engineer does not need to start from zero.

Evidence is more valuable than confidence

Technical incidents are full of convincing explanations that later turn out to be wrong.

That is normal.

Troubleshooting involves hypotheses.

The danger comes when a hypothesis is presented as fact.

“The firewall is blocking it” is very different from:

Connections succeed from one network segment but fail from another, name resolution is working and the service responds locally. Can you confirm whether traffic from this source is reaching the firewall and whether it is being denied?

The second version does something useful.

It separates evidence from interpretation.

It also gives the next person a specific technical question to answer.

I increasingly prefer escalations that say what we know, what we think and what we need to verify.

Those are three different things.

Keeping them separate makes collaboration much easier.

Timing depends on impact as well as knowledge

There is no universal rule for how long somebody should troubleshoot before escalating.

Twenty minutes may be too long in one incident and far too short in another.

Impact matters.

A fault affecting one non-critical endpoint can justify a slower investigation.

A fault affecting a critical service may require several people or suppliers to be engaged almost immediately while troubleshooting continues in parallel.

Risk matters too.

If the next troubleshooting step could make the situation worse, that may be the point to involve somebody with deeper knowledge.

Access matters.

There is little value in spending an hour proving something that another team can confirm directly from a system you cannot see.

And expertise matters.

Knowing the limits of your own knowledge is not weakness.

It is part of engineering judgement.

The escalation should contain a request

One of the most useful changes is to stop ending escalations with:

Can you have a look?

Sometimes that is genuinely all that can be asked.

Usually it can be more precise.

Can you confirm whether the service received these requests?

Can you check the authentication logs around this time?

Can you confirm whether any configuration changed before the fault began?

Can you validate the behaviour from the application side?

Can you advise whether this is expected behaviour?

Can you approve the next recovery step?

A clear request does two things.

It helps the recipient understand why they are involved.

And it tests whether the escalation has reached the right person.

If nobody can define what they need from the next team, the investigation may not yet be ready to leave the current one.

Escalation also carries responsibility

Passing an incident to another team does not necessarily mean ownership disappears.

This becomes especially important when a service crosses multiple systems.

A network team may need information from an application supplier.

An application team may need identity logs.

A cloud platform may depend on internet connectivity.

A vendor may need evidence from internal monitoring.

Each team can investigate its own component and still leave the overall service broken.

Someone needs to retain the end-to-end question:

Has the user’s service actually been restored?

That is why I see escalation as part of operational ownership rather than a mechanism for transferring a ticket elsewhere.

The purpose is not to make the problem somebody else’s.

The purpose is to involve the right capability while preserving the thread of the investigation.

Good communication is technical competence

Technical ability is sometimes described as though it exists separately from communication.

In complex environments, that distinction becomes difficult to maintain.

An engineer may understand a problem extremely well.

If they cannot explain the evidence, another team cannot use that understanding.

A specialist may know exactly what they need.

If nobody gives them the relevant timestamps, symptoms or test results, their expertise is harder to apply.

Good escalation therefore sits at the intersection of troubleshooting and communication.

It requires enough technical understanding to know what matters and enough clarity to explain it accurately.

The best escalation lets the next person begin where you stopped.

That saves time, avoids duplicated work and makes it more likely that the investigation follows evidence rather than assumptions.

Escalation is not the point where technical work ends.

Done properly, it is part of the technical work.

About the publication

I turn complex infrastructure and cybersecurity responsibility into resilient services, controlled change and evidence leaders can trust.