AI Can Explain the Runbook. It Cannot Own the Risk.
- zachary young

- Jul 22
- 5 min read
AI is useful until the moment someone has to own the decision.
Not a theoretical decision. Not a clean prompt. Not a tidy comparison between option A and option B.
A real one.
Traffic is under pressure. Users are feeling it. A DDoS event is in play. The primary path is no longer something you fully trust. There is a secondary path available, but moving production traffic is not free. It carries its own risk. It may solve the immediate problem. It may also expose something else.
That is when the fantasy version of AI runs out of road.
AI can explain what a DDoS is. It can describe common mitigation strategies. It can help draft an incident update. It can summarize a runbook, clean up a timeline, or give you a checklist of things to consider before a failover.
All of that can be helpful.
But it cannot own the call.
It cannot look at the room and understand the pressure. It cannot weigh the business context that never made it into the documentation. It cannot know which system is technically redundant but politically fragile. It cannot understand that the secondary path has been tested, but not under this exact kind of stress. It cannot feel the difference between waiting too long and moving too early.
When production gets real, the work is not just knowing the answer.
The work is deciding what risk you are willing to accept.
The Gap Between Advice and Accountability
A lot of AI conversations blur the line between advice and accountability.
That blur is dangerous.
Advice is cheap. Accountability is expensive.
Advice says, here are your options.
Accountability says, we are going to do this now, and I understand what might happen if it goes wrong.
During an incident, AI can help you organize the advice side. It can remind you to confirm impact, check monitoring, notify stakeholders, validate dependencies, document timestamps, and establish rollback criteria. That is genuinely useful, especially when adrenaline is high and everyone is moving fast.
But the hard part is not asking whether a failover might help.
The hard part is deciding whether this is the moment to invoke the DR posture you planned for on a much calmer day.
That decision belongs to people.
It belongs to the people who understand the architecture, the business impact, the customer pressure, the contractual commitments, the internal appetite for downtime, and the consequences of being wrong in either direction.
AI can support that conversation. It cannot replace it.
DR Is a Policy Question Before It Is a Technical Question
Disaster recovery gets talked about like it is mostly a technical problem.
It is not.
There are technical pieces, obviously. Secondary paths. Redundant services. Recovery objectives. Monitoring. Runbooks. Vendor contacts. Routing plans. Access. Testing. Rollback steps.
But the deepest DR questions are policy questions.
Who is allowed to declare an incident?
Who is allowed to initiate failover?
What impact threshold justifies moving traffic?
What is the expected recovery time?
What services are prioritized first?
Who communicates to the business?
Who communicates to customers?
What evidence do we need before we act?
How long are we willing to wait while the primary path is degraded?
What does success look like after the move?
What would make us roll back?
Those questions should not be invented in the middle of a DDoS event.
The incident will already be messy enough.
If the policy work is not done ahead of time, then the technical team ends up carrying not only the outage, but also the unresolved business decision. That is a heavy place to be.
A runbook can tell you how to move.
Policy tells you when you are allowed to move.
AI is much better at the first one than the second.
The Runbook Is Not the Decision
This is where people sometimes get fooled.
A runbook can feel like certainty. It has steps. It has headings. It has commands, contacts, systems, screenshots, and escalation paths. It makes the incident look solvable, because the page has an order to it.
But a runbook is not the decision.
A runbook is a tool for executing a decision that still has to be made.
The same is true for AI output. AI can produce an impressive checklist in seconds. It can say, assess impact, communicate with stakeholders, validate the alternate route, monitor latency, confirm DNS or routing behavior, and prepare rollback.
Great.
Now someone still has to decide.
Do we move now?
Do we hold?
Do we partially shift?
Do we wait for the provider?
Do we accept degraded performance while we gather more evidence?
Do we take the risk of changing the network path while customers are already impacted?
That is not a prompt problem.
That is an ownership problem.
The Human Context Is Usually the Missing Context
AI only knows what you give it.
Incidents are full of things nobody had time to put in the prompt.
A vendor relationship. A known weak spot. A previous test that passed technically but left people uneasy. A dependency that is not documented well. A business leader who cares about one workflow more than all the dashboards suggest. A customer-facing promise that changes the tolerance for waiting.
There is also the human layer.
Who is calm? Who is guessing? Who has seen this before? Who is giving confident answers because they are certain, and who is giving confident answers because they are nervous?
AI cannot read that room.
It cannot know when the team needs one clear owner instead of five smart suggestions. It cannot tell when people are avoiding the call because the downside is visible and the upside is only that things stop getting worse.
That is where judgment lives.
Not in the clean version of the problem.
In the messy version.
What AI Can Still Do Well
None of this means AI is useless during an incident.
It can be useful in the right lane.
It can help draft internal updates so the person leading the incident is not also fighting blank-page syndrome. It can turn notes into a timeline. It can summarize what changed between two status updates. It can help write a post-incident review. It can pressure-test a runbook before the incident ever happens. It can help teams create tabletop scenarios and ask uncomfortable questions about authority, communication, dependencies, and rollback.
It can also help after the fact.
What did we learn?
Where did the runbook help?
Where did the runbook assume too much?
Where did policy slow us down?
Where did ownership get fuzzy?
What decision did we make under pressure that should become a documented standard next time?
That is where AI can add a lot of value.
Not as the person making the call, but as a tool for turning experience into better preparation.
Better DR Starts Before the Incident
The real lesson is not that AI fails when things get serious.
The lesson is that organizations fail when they expect tools to compensate for unclear ownership.
A good DR posture should make the hard parts less improvised.
That does not mean every incident becomes easy. It means the team knows who can declare the event, who can approve the failover, what thresholds matter, who gets notified, what order systems come back in, and what tradeoffs the business has already accepted.
The technical plan matters.
The decision model matters just as much.
If the decision model is missing, the team will build it live, under stress, with partial information, while users are waiting.
That is not resilience.
That is hope with a runbook attached.
The Takeaway
AI can explain the runbook.
It cannot own the risk.
That distinction matters more as organizations put AI into more operational workflows. The closer AI gets to production, security, network paths, customer impact, and business continuity, the more important it becomes to know where assistance stops and accountability starts.
Use AI to prepare better. Use it to document better. Use it to challenge assumptions before the incident. Use it to clean up the mess afterward.
But when the rubber meets the road, someone still has to make the call.
Someone has to own the risk.
And that someone is not the model.



Comments