Skip to content
Article

Business Continuity When Third-Party Outages Are the New Normal

Two-thirds of outages now start at third parties. A guide to dependency mapping, vendor RTOs and fail-small design for business continuity.

Oğuzhan Gerçek··7 min read
Business Continuity When Third-Party Outages Are the New Normal

Short answer: If your business continuity plan is still written around the scenario of "our own data center goes down", you are blind to most of today's risk. According to the Uptime Institute's 2026 Annual Outage Analysis, roughly two-thirds of publicly reported outages involve third-party providers such as cloud, telecom and colocation companies. The first move is not buying another tool: it is mapping, layer by layer, which external services your critical business processes depend on, writing each provider's realistic recovery time into your plan, and designing a "stay up without the provider" mode for the flows that matter most.

The outage now comes from outside

Enterprise business continuity was built around scenarios where the generator fails to start or a storage array dies. Those scenarios are still valid, but the statistics point elsewhere. The Uptime Institute reports that per-site outage rates have been falling for five consecutive years, while roughly two-thirds of publicly reported outages involve cloud, telecom and colocation providers. In other words, IT teams are running their own infrastructure better; the outage has moved into the supply chain.

The cost side is not shrinking either. In the same research, 57 percent of respondents said their most recent major outage cost more than 100,000 dollars, and one in five reported losses above 1 million dollars. For a sense of tempo, ThousandEyes' network health tracking logged 649 global network outage events in the single week of August 24-30, 2026, a 22 percent increase over the week before.

Three lessons from the last twelve months

Looked at individually, every major outage has a different cause. Put them side by side and a common pattern emerges: a small internal change turns into a global business continuity event because of the provider's scale.

  • AWS, October 20, 2025. A race condition in DynamoDB's automated DNS management left the regional endpoint in us-east-1 unreachable. According to ThousandEyes' analysis, full recovery stretched into the next day and took more than 15 hours; global services such as Slack and Atlassian were affected. AWS's official post-event summary confirms the defect was a latent condition that had been sleeping in the code for years.
  • Microsoft Azure, October 29, 2025. An inadvertent configuration change in Azure Front Door triggered a global disruption affecting many services, including Microsoft 365. Microsoft has since published the lessons learned and tightened its deployment safeguards.
  • Cloudflare, November 18, 2025. A database permissions change inflated the Bot Management configuration file past its limit and crashed the proxy software. By Cloudflare's own report, the first errors appeared at 11:28 UTC and all services only returned to normal at 17:06 UTC.

None of these three events involved an attacker or a hardware failure. All three were routine changes propagating with a global blast radius. An IT team sitting on the customer side has zero chance of preventing these events; limiting their impact, however, is entirely in its own hands.

Who pays the bill: the CrowdStrike lesson

The clearest illustration of the financial side of a third-party outage is still the CrowdStrike update failure of July 2024. By Parametrix's estimate, the incident caused 5.4 billion dollars in direct losses for Fortune 500 companies alone, of which only 540 million to 1.08 billion dollars was covered by insurance. Most of the loss stayed on the companies' own books.

The lesson here is contractual: an SLA credit is not a compensation mechanism, it is a symbolic refund. Getting part of a monthly invoice back does not cover a stopped production line or a payment system that cannot process transactions. A business continuity budget should be built on the real business cost of an hour of downtime, not on SLA percentages.

Regulators are pointing at the same thing: DORA and Law 7545

Regulators saw this picture before most of us did. In the European Union, DORA, applicable since January 17, 2025, obliges financial entities to manage ICT third-party risk, maintain a register of ICT contracts, and prepare exit strategies for critical providers. Turkish finance and technology companies serving EU customers are links in that chain.

In Türkiye, an important step followed Cybersecurity Law No. 7545, which entered into force in March 2025: with the Cybersecurity Board decision of May 5, 2026, 15 critical infrastructure sectors were designated, including energy, finance, healthcare, transportation and digital infrastructure. Organizations in those sectors face an obligation to procure cybersecurity products and services from authorized providers, with administrative fines ranging from 1 million to 10 million TL. We previously examined the security and liability side of the supply chain from a KVKK perspective; the continuity side covered in this article is the second layer of the same map.

How to build the dependency map

The first step of third-party resilience is an inventory, and most organizations do not have one. The method works in this order:

  1. Start from the business process, not the technology. List 10-15 critical business flows such as "order intake", "payment collection" and "production planning".
  2. Write down the applications under each flow. ERP, CRM, email, the contact center platform: this is the visible layer.
  3. Write down the external services under each application. Cloud region, CDN, DNS provider, identity service (SSO/MFA), payment gateway, SMS and email delivery services, license servers. This invisible layer is where the real surprises live during an outage.
  4. Run the common-point analysis. How many critical flows pass through the same cloud region, the same DNS provider, the same identity service? This analysis shows how many processes stop at once when a single provider fails.
  5. Assign a vendor RTO to every dependency. Base it on real recovery times from actual incidents, not the uptime percentage on the provider's marketing page. The 15 hours in the AWS example teaches more than any "99.99%" claim.

Once this map exists, your target RPO and RTO values deserve a fresh look; see our RPO and RTO calculation guide for the method.

Fail small: an architecture for keeping outages contained

After its incidents, Cloudflare launched a resilience program it calls "Fail Small": staged rollout of configuration changes, interfaces between services designed on the assumption that the other side will fail, and break-glass procedures free of circular dependencies. The principles the provider set for itself apply one-to-one to the organizations consuming its services:

  • Define graceful degradation modes. When bot protection fails, should the site go down entirely, or should protection switch off while traffic keeps flowing? That decision must be made at design time, not in the middle of an outage.
  • Keep a static fallback. Holding a static copy of critical customer-facing pages with a different provider is the difference between a full outage and "slow but standing".
  • Avoid locking DNS and CDN into a single provider. Secondary DNS is low-cost insurance; a multi-CDN architecture is worth evaluating for critical traffic.
  • Test the offline scenario of your identity service. If your SSO provider sits in front of every system you access, the existence and tested state of break-glass accounts is vital.
  • Add vendor outages to your exercises. Disaster recovery drills usually simulate the loss of your own infrastructure; scenarios like "cloud region unreachable" or "SaaS CRM down for 8 hours" belong on the list. For the overall strategy, our enterprise disaster recovery strategy article is a good starting point.

Conclusion: start this week

Third-party resilience is not a purchasing decision; it is a discipline of visibility and design. Concrete steps you can take this week: build the dependency map for your five most critical business flows, compile each external provider's real outage history for the last two years, compare the compensation clauses in your SLAs against your actual business cost, and add at least one vendor-outage scenario to your next business continuity exercise. These four steps photograph the risk without requiring any budget, and that photograph then sets the priority of the architectural investments that follow. At eclit, this is exactly the order in which we approach business continuity planning for our managed services customers: visibility first, then design, and tools last.