Skip to content
Article

Business Continuity for SaaS Outages and Third-Party Risk

Salesforce, AI platforms and a nationwide ISP failure teach the same lesson. Your business continuity plan must now cover provider outages too.

Oğuzhan Gerçek··8 min read
Business Continuity for SaaS Outages and Third-Party Risk

Short answer: If your business continuity plan still covers only your own servers, your own data center and your own backups, it is not looking where outages actually come from anymore. Nine years of Uptime Institute data show that roughly two thirds of publicly reported outages originate with third-party providers such as cloud, telecom and colocation companies. The answer is not to flee SaaS; it is to map your dependencies, define a degraded mode of operation for every critical service, and test those scenarios exactly the way you test disaster recovery.

Three outages in one month, on three different layers

On the morning of September 16, 2026, Salesforce suffered access problems across all three of its operating regions. The outage began at 07:50 UTC, and the source of the problem took roughly three hours to diagnose. Early findings pointed to requests stalling while waiting on an internal login service. One of the world's largest SaaS companies, on day two of its own Dreamforce conference, watched its customers fail to open their CRM screens.

Two weeks earlier, on the morning of September 3, ChatGPT, Claude and Grok went down in the same window. Grok's incident was traced to a failure at a Memphis compute facility, and the statement mentioned impact on "compute partners". Companies wiring AI into their workflows discovered a new dependency layer that morning: the GPU infrastructure behind the model provider.

In Türkiye, on July 31, 2026, TürkNet subscribers experienced a nationwide access failure lasting hours. The company attributed the problem to an unexpected condition in its traffic routing systems; access only returned gradually at 16:16. The telling detail: customer support channels were unreachable during the outage as well. When your provider goes down, your channel for reaching that provider can go down with it.

Three incidents, three different layers: application (SaaS), AI infrastructure and the access network. The common thread is that none of the failures happened in your machine room. Yet the business that stopped that day was yours.

Outages moved outside, the bill stayed inside

This is not an impression; it is a measured trend. The Uptime Institute's 2026 Annual Outage Analysis reports per-site outage rates falling for the fifth consecutive year, while third-party IT and data center providers account for roughly two thirds of publicly reported outages over the nine-year tracking period. Individual facilities are being run better; but as dependency chains grow longer, a single provider failure spreads across a much wider area.

The bill is growing too. In the same analysis, 57 percent of respondents say their most recent major outage cost more than 100,000 dollars, and one in five puts the cost above 1 million dollars. The 2026 Hidden Costs of Downtime study by Oxford Economics and Splunk shows the picture at the top end of the market: unplanned downtime costs Global 2000 companies a combined 600 billion dollars a year, an average of 300 million dollars per company, and that cost has risen 50 percent in two years. A single incident is followed by an average 3.4 percent drop in share price.

Most of these figures come from global samples; no aggregate downtime cost data is published for Türkiye specifically. But the mechanics are identical everywhere: the e-commerce site cannot sell on campaign morning, the call center cannot see the CRM, the field team cannot receive work orders. The provider owns the outage; you own the damage. And the SLA credit in your contract typically refunds a portion of the monthly fee, not the lost revenue.

Why your DR plan does not cover SaaS

Classic disaster recovery is built on infrastructure you control: the primary site fails, the secondary site takes over. RPO and RTO targets are in your hands through replication and failover mechanisms. In SaaS, none of those levers are yours. When Salesforce goes down, you cannot fail over to a "standby Salesforce"; you refresh the provider's status page and wait.

The AWS us-east-1 outage of October 20, 2025 was the textbook case of this asymmetry. AWS's own post-event summary describes how a race condition in DynamoDB's automated DNS management wiped all IP records for the regional endpoint. According to ThousandEyes' analysis, the impact stretched past 15 hours from first packet loss to full recovery and spread to a long list of services including Slack, Atlassian and Snapchat. A large share of the companies that went down that day had no direct contract with AWS at all; they were customers of a SaaS that ran on AWS. This is concentration risk: when your different vendors depend on the same underlying layer, a single point of failure hides beneath the appearance of diversity.

This is the resilience version of the liability model we know from supply chain security. We covered why responsibility for a data breach cannot be outsourced to a vendor in our article on KVKK and supply chain liability; the same principle applies to downtime. Regulators are looking in the same direction. In the European Union, DORA has required financial entities to manage ICT third-party risk since January 2025, and Article 28 makes documented exit strategies for critical ICT providers mandatory. In Türkiye, the BDDK's Regulation on Banks' Information Systems and Electronic Banking Services imposes similar business continuity and audit obligations on outsourced services. It would be no surprise if the questions asked of banks today are asked of other sectors tomorrow.

Dependency mapping: the first and cheapest step

The first step in third-party resilience is not technology; it is inventory. Most organizations have no single map showing which business process depends on which SaaS, which cloud that SaaS runs on, and which identity provider sits in front of it. Without that map you can see neither your concentration risk nor which outage is actually critical.

In practice we start with three questions:

  • If this service goes down, which business process stops, and within how many minutes? Does revenue stop, does productivity dip, or does nobody notice?
  • What sub-dependencies sit behind this service? Which cloud region, which identity provider, which payment infrastructure?
  • Do we hold a current copy of the data in this service? If so, when was it last exported?

The answers to those three questions produce your criticality tiers on their own. Then you set a target per tier: if tier-one services carry a continuity expectation measured in minutes rather than hours, you compare the cost of staying single-provider against the cost of an alternative channel, explicitly. The method is the same business-cost approach we describe in our RPO and RTO calculation guide; this time you are applying the targets to your vendors instead of your own infrastructure.

Degraded mode: how work continues during an outage

The second step is a written answer, per critical service, to the question "how does work continue while this service is gone". We call this a degraded mode of operation, and it has three components:

  1. An independent copy on the data side. Regular exports of your SaaS data should live in an environment you control. An export that can show the day's customer appointments when the CRM is unreachable turns an outage from a crisis into an inconvenience for most sales teams.
  2. A temporary procedure on the process side. If orders cannot be entered into the SaaS, where are they recorded, and who replays them in what sequence once the outage ends? If that procedure is not written down, it gets invented mid-outage, and that usually ends in data loss.
  3. An independent channel on the communication side. As the TürkNet case reminded everyone, you need a communication path that does not depend on the provider or the affected infrastructure. If the channel you would use to talk to your team during a crisis runs on top of the failing service, your plan exists only on paper.

And above all of it: none of these scenarios count as a plan until they are tested. Add "our most critical SaaS vendor is unreachable for 8 hours" to your tabletop exercise. We described what happens to organizations that never restore their backups in our piece on testing versus paper plans; the same rule holds for third-party scenarios. An untested degraded mode is not a mode of operation.

Three steps to take this week

The part of this work that needs no budget approval fits in a week:

  1. List your ten most critical external dependencies, including SaaS, cloud, connectivity and identity providers. For each one, write the answer to "what happens if it goes down" in a single sentence.
  2. For every service on that list, put the SLA commitment and your real expectation side by side. The gap between them is risk you are accepting; it should be accepted knowingly.
  3. Add at least one third-party outage scenario to your next business continuity exercise, and share the result with management.

As a managed services provider, a significant part of our day goes into watching these dependency chains alongside our customers' own infrastructure. Our experience is this: you cannot prevent a provider outage, but you can prevent it from catching you unprepared. That difference is the difference between staring at a status page that morning and keeping the business running.