Every reliability program you have run over the past five years has been working. Uptime Institute’s 2026 outage analysis records a fifth consecutive year of falling per-site outage frequency.
However, the report notes the pace of improvement has slowed, and roughly one in ten operators still describe their last outage as seriously or severely damaging.
Five years of improvement came from failover drills and DR planning inside environments you operate. The services your providers run sat outside that work, and that is where your remaining exposure lives.
Resiliency assessments concentrate on the systems you operate
Uptime finds that resiliency assessments remain focused on internal systems rather than on external and systemic risks. The audit that produced five years of improvement is the same audit that never opens the category, still causing serious harm.
Gartner reached a parallel conclusion from the risk side. In its risk survey of 294 risk executives, third-party viability topped the list for two consecutive quarters and cloud concentration entered the top five at 62%, with a warning that many organizations would face severe disruption if a single provider failed.
Regulators have since formalized the exposure. EU DORA has applied to roughly 22,000 financial entities since January 2025, mandating concentration risk assessment across third-party providers, a register of ICT dependencies, and contractual exit provisions.
What can you do?
Power failures remain the leading cause of impactful outages, and your review covers that category well.
Before your next assessment signs off, list the external providers each revenue-critical flow cannot run without. That list is usually nowhere on paper.
Stay updated with Simform’s weekly insights.
External dependency failures run long because the provider controls recovery
Uptime finds external infrastructure failures growing more prominent in publicly reported outages, with fiber and connectivity incidents both rising and more likely to turn into extended disruptions.
When something you operate fails, recovery runs on your own team. When a connectivity provider or a third-party API fails, recovery runs on the provider’s schedule, and your monitoring surfaces the symptom long before the cause.
IBM’s breach research measures a security-incident lifecycle, and within that scope, third-party and supply-chain compromises took the longest of any attack vector to resolve at 267 days.
The mechanism transfers even where the number does not. Incidents crossing a trust boundary run long because internal instrumentation is not watching that boundary.
What can you do?
Microsoft’s Well-Architected reliability monitoring guidance names the fix. Track retry behavior and transient fault rates against your dependencies, and a degrading external service shows up as your alert while the customer still has a working product.
Instrument the paths you do not own the way you instrument the ones you do. If your observability spend already feels heavy, redirect coverage toward the trust boundary before adding tools.
Vendor fees set the ceiling on your SLA credit
Uptime puts 57% of operators above $100,000 for their most recent major outage, and for the second year running, one in five above $1 million. The bill lands on your side of the contract regardless of who caused the failure.
A vendor SLA credit is calculated against the fees you pay that vendor. The revenue you lose while their service is down never enters that math, so the payout and the damage sit on unrelated scales.
Microsoft concedes the point in its own mission-critical guidance. Even where a third-party routing service advertises a 100% SLA, the framework notes that financial reparation carries little weight once outage impact is large. It tells architects to build redundancy regardless of the contract number.
What can you do?
The gap between credits and actual loss is yours to calculate and nobody else will hand you the figure. Put a revenue-per-hour number beside each critical external dependency and compare it to the credit its SLA would pay.
Dependency mapping in Azure starts with failure mode analysis
Microsoft’s failure mode analysis guidance tells teams to decompose critical flows and identify dependencies both internal and external to the workload, capturing availability SLAs and scaling limits for each.
Its reliability testing guidance sets the priority order, starting with critical-flow dependencies that have no fallback path and testing partial and cascading failures rather than only complete outages.
Neither runs by default. Both are configuration decisions someone has to own.
| Dependency tier | Expected behavior on failure | What to test |
| No fallback exists | Flow stops, customer sees it | Whether anyone is paged before the customer notices |
| Degraded mode designed | Flow continues, reduced function | Whether degradation triggers automatically |
| Redundant provider configured | Traffic reroutes | Whether failover fires, and how long cutover takes |
Record each external provider’s SLA, the tier it belongs to, and the fallback you have tested. That is an audit with a completion date rather than a standing program.
From the field
We built and validated a recovery framework on this sequencing for an investment firm on Azure, where DR testing ran 60% faster and the resulting posture held under FDIC and PCI DSS audit.
Documenting a fallback and confirming it fires are different pieces of work, and only the second one survives an incident.
Model APIs are the newest external dependency in your critical path
Most AI features shipping this year run on model endpoints your team does not operate. Microsoft’s Well-Architected AI architecture guidance is direct about the exposure, noting that knowledge and tool layers frequently depend on external systems that introduce delays or availability issues affecting response quality.
Its application design guidance names the mitigation, which is an AI gateway that routes requests to alternate providers when the primary is unavailable. Multi-provider failover is a design decision.
McKinsey describes the same architecture from the buyer’s side, characterizing enterprise generative AI as accessing a foundation model through an API and warning that agents embedded in core operations have to weigh reliance on externally hosted endpoints.
Your payment provider has been a tolerated dependency for years, absorbed quietly because it sits at the edge of the product. Model endpoints have sat in the core value path since the first release, so the first serious outage reaches a customer.
Buying the capability was the right call, and it put external dependency in the core value path at a headcount that has nobody assigned to map it.
Of the outside dependencies you would be down with right now, how many has anyone written a recovery plan for?
Simform runs Azure environments as an engineering discipline, backed by Azure Expert MSP status. See what that looks like in practice through engineering-led Azure managed services.