Post ·
Blast Radius Is a Query, Not a Diagram
Blast radius is a walk over CMDB relationships, not a picture, and only as honest as its paths through services. Three checks before an agent trusts it.


Every CMDB has a dependency diagram somewhere. It hangs in a wiki, it has a legend, and somebody was proud of it once. Ask it the only question that matters during an outage — this host just died, who do we tell? — and it goes quiet. It was never going to answer. A diagram is a picture of what someone drew. Blast radius is a question about what you modeled.
Blast radius is a walk: start at the element that failed, follow its relationships toward the business, count the hops, and stop where a customer or a process would notice. The picture is a souvenir of the walk. The walk is the thing.
This used to be an awkward bridge call when the answer was wrong. Now the reader is increasingly an agent deciding whether a change is safe, and the questions being asked about AI agents and the CMDB keep landing on the same one: how far does a failure spread? A wrong answer used to cost a meeting. It now costs whatever the agent touches next.
What the walk actually does
In the Common Service Data Model (CSDM), the walk moves upward through the layers, one relationship at a time:
start the failed element (host, cluster, network)
hop 1 what runs on it (applications, databases)
hop 2 the services that depend on it (production, not test)
hop 3 what the services serve (business application, offering)
hop 4 what the business calls it (business service, capability, process)
stop at something with an owner and a promiseNothing in that walk looks at coordinates, colors or which box is drawn above which. It reads relationship types and their direction. Two models can render identically and walk completely differently, which is the polite way of saying the diagram was never the source of truth.
Three models, three answers
Blueprint Modeler ships with application examples, and I've published three of them as reference diagrams. Each one answers the blast-radius question differently, and each answer comes from the shape, not the drawing.
- Online store checkout. In the checkout model, when db-prod-01 fails, the walk reaches the orders database, the production service, the Checkout application and the North America offering, then the capability and the business service: six elements in four hops. When web-test-01 fails, it never touches a customer-facing offering. Same application, two environments, two very different afternoons — because production and test are separate services, not tags.
- HR self-service portal. In the HR portal model, when the production service fails, the walk reaches six elements in two hops, and one of them is the business process “Onboard a new hire.” The first sentence of the incident writes itself: onboarding is affected. That is what a process anchor buys you.
- Shared database platform. In the database platform model, when one PostgreSQL node fails, the walk reaches eight of eight elements in four hops. Every service on the cluster, both applications, the platform's offering. If the cluster is built to survive a node, the model hasn't been told, and the walk believes the model.
Where the walk lies
A blast radius is only as honest as its relationships. These are the ways it goes wrong, roughly in order of how often:
- The missing middle. A business application pointed straight at a server skips the service, so the walk jumps from one host to the whole application — every environment, every region. I wrote about why business apps shouldn't point at servers; this is what it costs at 2 a.m.
- Relationships pointing the wrong way. If a dependency was stored backward, the walk stops at the first hop and reports that nothing upstream is affected. A blast radius of zero is not good news. It is usually a data-entry decision from three years ago.
- Redundancy that isn't modeled. Two nodes behind a cluster with no cluster record means every node failure reads as a full outage. The walk isn't wrong; it is faithfully reporting a single point of failure that someone forgot to draw as two.
- Services outside any offering. The walk reaches a service and stops, because no offering contains it. Nobody owns telling anyone. In the database example, the Reporting service consumes the platform without being in its contract, so the maintenance calendar has no reason to mention it.
- Relationships that outlived the thing. Decommissioned hosts still wired to live services make the radius bigger than reality. Over-alerting is how teams learn to stop reading alerts.
A diagram that is wrong looks exactly like a diagram that is right. A walk that is wrong usually tells you, if you ask it about a failure you already understand.
Before an agent trusts it
If an agent is going to read blast radius before it acts — and it will — three questions decide whether the answer deserves to be believed:
- Does every path from infrastructure to the business go through a service? If not, the walk over-reports in some places and under-reports in others, and you can't tell which.
- Does every walk end at something with an owner? An offering, a process or a capability with a named person behind it. A walk that ends at an orphaned service has found the problem and nobody to give it to.
- Has the walk been checked against a real incident? Take the last three outages, run the walk from the failed element, and compare it with who actually got paged. The gaps are your backlog. Until then, the agent reads and proposes; it doesn't act. Absent, not disabled is still the right default for write tools.
Copy this: the check
CHECK: can this blast radius be trusted?
FOR each host, cluster or network element E (operational):
walk = follow relationships upward from E
(runs on, depends on, contains, uses, provides)
IF walk reaches a business application without passing an
application service: flag "missing middle"
IF walk ends at a service with no offering:
flag "no owner to tell"
IF walk is empty but E has dependents in discovery:
flag "direction or gap"
IF E is one node of a redundant pair and the walk reaches
every dependent: flag "redundancy not modeled"
CALIBRATE: replay the last three incidents; compare the walk
with who was actually paged
REPORT flags by business criticality - fix production firstTry it yourself
Open Blueprint Modeler, load any of the three examples, right-click an element and choose Show blast radius. Then delete one relationship and run it again. Watching the radius shrink by a whole offering because one line went missing is the fastest argument for CSDM governance I know, and it takes less time than the meeting about it.