Summarize with AI

Not enough time? get the key points instantly.

Every reliability program you have run over the past five years has been working. Uptime Institute’s 2026 outage analysis records a fifth consecutive year of falling per-site outage frequency.

However, the report notes the pace of improvement has slowed, and roughly one in ten operators still describe their last outage as seriously or severely damaging.

Five years of improvement came from failover drills and DR planning inside environments you operate. The services your providers run sat outside that work, and that is where your remaining exposure lives.

Resiliency assessments concentrate on the systems you operate

Uptime finds that resiliency assessments remain focused on internal systems rather than on external and systemic risks. The audit that produced five years of improvement is the same audit that never opens the category, still causing serious harm.

Gartner reached a parallel conclusion from the risk side. In its risk survey of 294 risk executives, third-party viability topped the list for two consecutive quarters and cloud concentration entered the top five at 62%, with a warning that many organizations would face severe disruption if a single provider failed.

Regulators have since formalized the exposure. EU DORA has applied to roughly 22,000 financial entities since January 2025, mandating concentration risk assessment across third-party providers, a register of ICT dependencies, and contractual exit provisions.

What can you do?

Power failures remain the leading cause of impactful outages, and your review covers that category well.
Before your next assessment signs off, list the external providers each revenue-critical flow cannot run without. That list is usually nowhere on paper.

Stay updated with Simform’s weekly insights.

External dependency failures run long because the provider controls recovery

Uptime finds external infrastructure failures growing more prominent in publicly reported outages, with fiber and connectivity incidents both rising and more likely to turn into extended disruptions.

When something you operate fails, recovery runs on your own team. When a connectivity provider or a third-party API fails, recovery runs on the provider’s schedule, and your monitoring surfaces the symptom long before the cause.

IBM’s breach research measures a security-incident lifecycle, and within that scope, third-party and supply-chain compromises took the longest of any attack vector to resolve at 267 days.

The mechanism transfers even where the number does not. Incidents crossing a trust boundary run long because internal instrumentation is not watching that boundary.

What can you do?

Microsoft’s Well-Architected reliability monitoring guidance names the fix. Track retry behavior and transient fault rates against your dependencies, and a degrading external service shows up as your alert while the customer still has a working product.

Instrument the paths you do not own the way you instrument the ones you do. If your observability spend already feels heavy, redirect coverage toward the trust boundary before adding tools.

Vendor fees set the ceiling on your SLA credit

Uptime puts 57% of operators above $100,000 for their most recent major outage, and for the second year running, one in five above $1 million. The bill lands on your side of the contract regardless of who caused the failure.

A vendor SLA credit is calculated against the fees you pay that vendor. The revenue you lose while their service is down never enters that math, so the payout and the damage sit on unrelated scales.

Microsoft concedes the point in its own mission-critical guidance. Even where a third-party routing service advertises a 100% SLA, the framework notes that financial reparation carries little weight once outage impact is large. It tells architects to build redundancy regardless of the contract number.

What can you do?

The gap between credits and actual loss is yours to calculate and nobody else will hand you the figure. Put a revenue-per-hour number beside each critical external dependency and compare it to the credit its SLA would pay.

Dependency mapping in Azure starts with failure mode analysis

Microsoft’s failure mode analysis guidance tells teams to decompose critical flows and identify dependencies both internal and external to the workload, capturing availability SLAs and scaling limits for each.

Its reliability testing guidance sets the priority order, starting with critical-flow dependencies that have no fallback path and testing partial and cascading failures rather than only complete outages.

Neither runs by default. Both are configuration decisions someone has to own.

Dependency tier Expected behavior on failure What to test
No fallback exists Flow stops, customer sees it Whether anyone is paged before the customer notices
Degraded mode designed Flow continues, reduced function Whether degradation triggers automatically
Redundant provider configured Traffic reroutes Whether failover fires, and how long cutover takes

 

Record each external provider’s SLA, the tier it belongs to, and the fallback you have tested. That is an audit with a completion date rather than a standing program.

From the field

We built and validated a recovery framework on this sequencing for an investment firm on Azure, where DR testing ran 60% faster and the resulting posture held under FDIC and PCI DSS audit.

Documenting a fallback and confirming it fires are different pieces of work, and only the second one survives an incident.

Model APIs are the newest external dependency in your critical path

Most AI features shipping this year run on model endpoints your team does not operate. Microsoft’s Well-Architected AI architecture guidance is direct about the exposure, noting that knowledge and tool layers frequently depend on external systems that introduce delays or availability issues affecting response quality.

Its application design guidance names the mitigation, which is an AI gateway that routes requests to alternate providers when the primary is unavailable. Multi-provider failover is a design decision.

McKinsey describes the same architecture from the buyer’s side, characterizing enterprise generative AI as accessing a foundation model through an API and warning that agents embedded in core operations have to weigh reliance on externally hosted endpoints.

Your payment provider has been a tolerated dependency for years, absorbed quietly because it sits at the edge of the product. Model endpoints have sat in the core value path since the first release, so the first serious outage reaches a customer.

Buying the capability was the right call, and it put external dependency in the core value path at a headcount that has nobody assigned to map it.

Of the outside dependencies you would be down with right now, how many has anyone written a recovery plan for?

Simform runs Azure environments as an engineering discipline, backed by Azure Expert MSP status. See what that looks like in practice through engineering-led Azure managed services.

Stay updated with Simform’s weekly insights.

Hiren is CTO at Simform with an extensive experience in helping enterprises and startups streamline their business performance through data-driven innovation.

Sign up for the free Newsletter

For exclusive strategies not found on the blog

Revisit consent button
How we use your personal information

We do not collect any information about users, except for the information contained in cookies. We store cookies on your device, including mobile device, as per your preferences set on our cookie consent manager. Cookies are used to make the website work as intended and to provide a more personalized web experience. By selecting ‘Required cookies only’, you are requesting Simform not to sell or share your personal information. However, you can choose to reject certain types of cookies, which may impact your experience of the website and the personalized experience we are able to offer. We use cookies to analyze the website traffic and differentiate between bots and real humans. We also disclose information about your use of our site with our social media, advertising and analytics partners. Additional details are available in our Privacy Policy.

Required cookies Always Active

These cookies are necessary for the website to function and cannot be turned off.

Optional cookies

Under the California Consumer Privacy Act, you may choose to opt-out of the optional cookies. These optional cookies include analytics cookies, performance and functionality cookies, and targeting cookies.

Analytics cookies

Analytics cookies help us understand the traffic source and user behavior, for example the pages they visit, how long they stay on a specific page, etc.

Performance cookies

Performance cookies collect information about how our website performs, for example,page responsiveness, loading times, and any technical issues encountered so that we can optimize the speed and performance of our website.

Targeting cookies

Targeting cookies enable us to build a profile of your interests and show you personalized ads. If you opt out, we will share your personal information to any third parties.