On one cloud-backup product I worked with, we kept seeing customer escalations involving missing backup files.
The individual cases were being handled.
A senior technical specialist investigated them. Customers received answers. Service was restored when necessary. Tickets were closed.
But after several similar incidents, I went back through the history and logs across cases.
What looked like a sequence of separate technical problems started to look like something else.
The same class of issue was entering the organization through different channels — tickets, chats, direct conversations — being investigated largely as an individual case, and then disappearing from management attention once the immediate customer problem had been resolved.
There was no single owner for the end-to-end process.
That was the point where I stopped treating recurrence as only a reliability problem.
A service can be restored while the underlying problem remains unresolved.
And when the same class of incident keeps returning, the useful question is no longer only:
What failed technically?
It is also:
What in the operating model allowed the same problem to survive the previous incident?
That distinction changed how I approached the problem — from handling escalations one by one to building a system for intake, classification, prioritization, RCA and follow-up.
The organization was solving cases, not managing the problem
The existing process was not obviously broken.
That was part of the difficulty.
Support escalated an issue.
A specialist investigated it.
Someone removed the immediate blocker.
The customer received an answer.
The case was closed.
Every participant could reasonably say that their part had been completed.
But no one owned the result across those boundaries.
That left several questions unanswered:
- Were these genuinely separate incidents or one recurring problem class?
- Did the root cause require product, engineering or infrastructure work?
- Should the risk be eliminated or consciously accepted?
- How should preventive work compete with roadmap capacity?
- Who remained accountable after the immediate service had been restored?
The organization had owners for activities.
It did not consistently have an owner for the outcome.
That difference matters because local completion can create the appearance of control while the systemic problem remains untouched.
The first intervention was operational, not technical
My first step was not an architecture initiative.
It was to create one controlled path for escalations.
I prepared a common model for working with them and aligned it with the relevant stakeholders.
Support began routing escalations through a shared board instead of allowing them to arrive through multiple informal channels.
Incidents were classified and clustered.
Priority became visible.
Technical specialists were involved when their expertise was actually required rather than acting as an informal intake point for every difficult case.
At the beginning, this required a significant amount of manual coordination from me.
I confirmed the problem class, set priority, involved technical expertise where necessary, moved critical issues into work and kept stakeholders informed.
That was useful as an emergency operating layer.
But it was not the target state.
The target was a process that could work without depending on one person continuously pushing it forward.
So the next step was to formalize the workflow, agree it with the team and stakeholders, and turn the temporary coordination mechanism into a repeatable operating process.
The transition was from:
a manager personally coordinating incidents
into:
a system that made intake, priority, ownership and follow-up explicit.
As the process matured, my role became closer to a facilitator or referee than a dispatcher.
That was an important test.
If a process works only because one manager remembers every issue, chases every dependency and keeps every stakeholder aligned personally, the organization has not really solved the operating problem.
It has created a human workaround.
The immediate result was not fewer incidents. It was better control
The first effect was predictability.
Escalations entered through a known path.
Similar cases could be seen together instead of living in separate conversations.
Priority became explicit.
Stakeholders could see which operational problems were consuming team capacity and what other work would be displaced when a new issue moved up the queue.
The process also became less dependent on my direct orchestration.
Senior technical expertise was used more selectively, where deeper investigation was actually required.
Over time, broader RCA allowed the team to address the highest-priority recurring problem classes.
Some causes were fixed.
Others remained as consciously accepted risks.
That distinction was important.
The goal was not to eliminate every possible incident.
The goal was to ensure that recurrence resulted in an explicit management decision rather than another isolated technical response.
RCA only matters when it changes a decision
A root-cause analysis can be technically correct and still create very little organizational value.
Finding the cause is only the first half of the work.
The more important question is what changes after the cause becomes known.
In this case, recurring problem classes only began to disappear when RCA was connected to prioritization and actual capacity decisions.
Not every root cause was eliminated.
Some risks were consciously accepted because addressing them would have consumed capacity that was more valuable elsewhere.
That is not necessarily a reliability failure.
Organizations always operate under constraints.
The management responsibility is to make the trade-off explicit.
The useful sequence became:
incident → recovery → RCA → business prioritization → fix or accepted risk → follow-up
Without the second half of that sequence, RCA becomes a record of the past rather than a mechanism for changing the future.
Priority is a business decision, not a severity label
Once incidents were visible together, prioritization became a management problem rather than a ticket-management problem.
The first rule was straightforward: issues blocking fundamental business functions came first.
After that, priority depended on several factors, including functional criticality, recurrence frequency and the affected user segment.
I owned that trade-off within the area.
This mattered because technical urgency and business priority are not identical.
A technically unpleasant problem can have limited business consequence.
A relatively small defect can repeatedly affect a strategically important group of users.
And a system can appear healthy while a known unresolved risk remains open.
I saw this directly after one website failure.
Normal operation was restored.
The service was working again.
But the underlying cause had not been removed, and we had reason to expect that the same failure could return within the following months.
Operationally, the incident was over.
Managerially, the problem was still open.
That distinction prevented “service restored” from becoming a false signal that no further decision was required.
Cross-functional boundaries exposed the real limit of the model
The hardest part came when the root cause sat outside my area of ownership.
The required work belonged to another technical function.
It was not especially large, but because the issue did not interfere with current operation, it repeatedly lost against that function’s local priorities.
Inside my area, the risk remained visible.
It had survived recovery, RCA and prioritization.
But once the solution crossed an organizational boundary, there was no equally strong mechanism for resolving whose priority should win.
Persistent escalation eventually substituted for the missing cross-functional decision mechanism.
It worked eventually.
It was not scalable.
That experience exposed a broader management problem.
A team can behave completely rationally inside its own backlog and still make the company-level outcome worse.
If each function optimizes only its local priorities, the business criticality of a problem can weaken every time ownership crosses a boundary.
The missing mechanism was explicit cross-functional prioritization: a place where business impact, local capacity and competing commitments could be compared rather than resolved through persistence.
Transparency changed the stakeholder conversation
Before the common board and classification model existed, there was no reliable portfolio view of incident work.
People remembered individual cases.
They did not have a consistent view of recurring problem classes, current operational demand or how much capacity was being consumed by interruptions.
Once that became visible, the management conversation changed.
A stakeholder could still say that a new problem was urgent.
But instead of simply accepting or rejecting that claim, I could show the work already consuming the team’s capacity and ask a different question:
Is this more important than the work it will displace?
If the answer was yes, the affected stakeholders were involved in the trade-off.
Regular group planning made those decisions easier because priority conflicts no longer had to be negotiated independently every time.
This did not remove urgency.
It made the economics of urgency visible.
Without that transparency, missed plans could always be explained by saying that unexpected problems kept appearing.
With it, every interruption became visible as a capacity-allocation decision.
That made the system more uncomfortable in one sense.
It also made it more honest.
Strong experts should not become hidden infrastructure
The senior technical specialist in the original process was valuable because some incidents genuinely required deep expertise.
The problem was not that the specialist was involved.
The problem was that expertise had also become part of the routing mechanism.
Difficult cases naturally found their way to the person most likely to solve them.
That works surprisingly well for a while.
It also hides process weakness.
A mature operating model separates expertise from coordination.
Support owns intake and initial classification.
The operating process makes priority and flow visible.
Technical specialists are involved when their judgment is necessary.
Problem history lives in a shared system rather than in someone’s memory.
That protects scarce expertise from becoming the hidden infrastructure holding an immature process together.
It also gives specialists more capacity to work on systematic causes rather than repeatedly reconstructing the history of individual escalations.
What recurrence taught me about management
The important lesson from that experience was not that every organization needs a particular incident process.
The process itself was contextual.
The management pattern was more general.
A repeated incident should trigger more than another technical investigation.
It should trigger questions about the system around the incident:
- Is there an owner for the end-to-end outcome?
- Can the organization distinguish recovery from problem resolution?
- Do RCA findings actually change priority, capacity or policy decisions?
- Can a risk be consciously accepted instead of simply remaining unfixed?
- Does business priority survive when work crosses functional boundaries?
- Can stakeholders see what gets displaced when something new becomes urgent?
- Is scarce technical expertise being used for diagnosis, or compensating for missing process ownership?
- Would the process continue to work if the manager currently coordinating it disappeared?
The executive implication is not that leaders should personally manage incidents.
My own experience pointed in the opposite direction.
Leadership value came from recognizing that recurring technical symptoms were exposing an operating-model defect, creating temporary control, turning that control into a repeatable mechanism, and then reducing the organization’s dependence on personal intervention.
A single incident may tell you what failed in the system.
A repeated incident often tells you what is failing in management.