What Is Uptime and How Is It Calculated? 99.9% vs 99.99%
Uptime is the share of time a system stays up. How to calculate it, how much downtime each target allows, and what to check in an SLA.

Short answer: Uptime is the share of a given period during which a system or service stays up: time up / total time × 100. 99.9% uptime allows about 8 hours 46 minutes of downtime a year, and 99.99% about 52.6 minutes. In a contract the percentage says little on its own; what counts as downtime, the period and the point it is measured from, and whether planned maintenance counts should all be written down too.
What is uptime?
Uptime measures how long a system or service stays up. AWS defines availability along the same lines: the time a workload is available for use, as a share of the total time measured.
Uptime and availability do not always tell the same story. One of the definitions in the NIST glossary describes availability as the property of being accessible and usable on demand by an authorized entity. An application whose server is up but which cannot reach its database looks fine from the machine's side; for users, it is unavailable.
The word has a second, server-side meaning: on Linux, the uptime command shows how long the system has been running, that is, the time since the last boot. That tells you about the machine; on its own, it says nothing about whether the service is available.
How is uptime calculated?
The basic formula, as AWS puts it, is availability = uptime / (uptime + downtime). Multiply the result by 100 to get a percentage.
An example: a 30-day month has 43,200 minutes. If two outages that month lasted 30 minutes in total, uptime is (43,200 − 30) / 43,200 = 99.93%. That meets a 99.9% target but falls short of 99.95%.
Google's SRE book describes a second method: request-based availability, the share of requests that succeed. In the book's example, a system serving 2.5 million requests a day with a daily 99.99% target can return at most 250 errors. According to the book, this measure is more useful than outage length for services whose load varies over the course of a day or week.
99%, 99.9%, 99.95% and 99.99%: allowed downtime
The calculation is simple: allowed downtime = (1 − target) × length of the period. The figures below use a 365-day year (525,600 minutes) and a 30-day month (43,200 minutes):
- 99%: 3.65 days a year (87.6 hours), 7.2 hours a month
- 99.9%: 8.76 hours a year (8 hours 45 minutes 36 seconds), 43.2 minutes a month
- 99.95%: 4.38 hours a year (4 hours 22 minutes 48 seconds), 21.6 minutes a month
- 99.99%: 52.56 minutes a year, 4.32 minutes a month
The availability table in Google's SRE book gives the same results. The book does not state the length of the year it uses, but its figures match a 365-day year and a 30-day month exactly. In a leap year, the 99.99% allowance rises to 52.7 minutes.
Each extra "nine" cuts the allowed downtime to a tenth: going from 99.9% to 99.99% takes the yearly allowance from 8.76 hours down to 52.6 minutes. Which month you measure matters too; the same 99.9% target allows 44.6 minutes in a 31-day month and 40.3 minutes in a 28-day February.
SLA vs SLO vs SLI
Google's SRE book separates the three terms like this:
- SLI (service level indicator): A carefully defined quantitative measure of some aspect of the level of service, such as the share of successful requests.
- SLO (service level objective): A target value or range for a service level measured by an SLI, such as "99.9% of requests will succeed".
- SLA (service level agreement): A contract that contains SLOs and sets out the consequences of meeting or missing them.
The book's test for telling them apart is simple: ask what happens if the SLO is missed. If there is no explicit consequence, you are almost certainly looking at an SLO. Nor should the target be 100%; the book says 100% is probably never the right reliability target. According to the SRE workbook, the error budget is 100% minus the SLO. For a 99.9% SLO that is 0.1%, which over the four-week window the workbook recommends comes to about 40 minutes in time-based terms.
What to check in an SLA's uptime clause
The uptime and SLA entries in our glossary make the same point: a percentage on its own is not enough. Cloud providers' SLAs are a good illustration of what to look for:
- What counts as downtime: In the AWS EC2 SLA, an individual instance is unavailable when it has no external connectivity. As long as the instance can be reached, errors from the application running on it do not count under that definition.
- Scope: The same SLA commits to 99.99% for deployments spread across two or more availability zones, and to 99.5% for a single instance.
- Short outages: In the Google Compute Engine SLA, outages shorter than a minute do not count toward downtime.
- Planned maintenance: In Oracle's cloud hosting and delivery policies, scheduled maintenance periods are excluded from the unplanned downtime calculation.
- Measurement window: These contracts calculate uptime monthly, so a bad week can be offset by the rest of the month.
- Compensation: At AWS and Google, the remedy for a breach is a service credit worth a percentage of the monthly bill. Google requires credit claims within 60 days, with logs showing the downtime; without your own measurements, a breach is hard to prove.
We have collected other questions for providers in 10 questions to ask when choosing a managed service provider.
How is uptime measured?
To measure the uptime users actually see, you test the service from the outside, the way a user would. The SRE book calls this black-box monitoring: testing externally visible behavior as a user would see it. In practice, that means synthetic checks:
- AWS CloudWatch Synthetics canaries follow the same routes as a customer and can run as often as once a minute, so the service is verified even when there is no traffic.
- Google Cloud uptime checks send requests from several locations around the world, and a check only fails when no response comes back from more than one location. A network problem at a single location does not turn into a false alarm.
The check interval matters too: a check that runs once a minute cannot see an outage that starts and ends between two checks. That is why synthetic checks should be read together with the application's own error and request metrics. We describe how this is set up on our 24/7 monitoring page.
MTTR and MTBF: what drives uptime?
AWS also writes the formula with two averages: availability = MTBF / (MTBF + MTTR). MTBF (mean time between failures) is the average time between a workload starting normal operation and its next failure; MTTR (mean time to repair) is the time until the failed part is repaired or returned to service.
The arithmetic shows that two routes lead to the same place. With an MTBF of 1,000 hours and an MTTR of 1 hour, availability is 99.90%. Cutting MTTR to half an hour and raising MTBF to 2,000 hours both take it to 99.95%. In AWS's words, failing less often and recovering faster both lead to higher availability.
Architecture is part of the equation. When a server behind a load balancer fails its health checks, it is taken out of service and requests go to the healthy servers. Dependencies, on the other hand, pull the number down. According to AWS, the theoretical maximum is the product of the availability of every component that must be working; for a service that depends on two 99.9% components, that ceiling is 99.8%. AWS stresses that this is a rough estimate, but also that fewer dependencies mean a lower likelihood of failure. We cover the preventive maintenance side in our system reliability service.
Uptime is not the same as Tier levels
The organization that runs the data center Tier classification is called Uptime Institute, but Tier levels are not uptime percentages. The classification focuses on data center infrastructure, and according to the institute's own explanation, the current standard does not assign availability predictions to Tier levels; references to expected downtime per year were removed in 2009. Running in a Tier III facility does not mean your application will hit any particular uptime percentage. We cover Tier levels in our data center article.
Frequently asked questions
What does uptime mean? The time a system or service stays up and running. It is usually given as a percentage of the total time in a given period.
How much downtime is 99.9% uptime? About 8 hours 46 minutes in a 365-day year, and 43.2 minutes in a 30-day month.
What is the difference between 99.9% and 99.99%? A 99.99% target allows a tenth of the downtime that 99.9% does: 52.6 minutes a year instead of 8.76 hours.
Are uptime and availability the same thing? They are close but not identical. Uptime says the system is running; availability says users can actually use the service.
Is 100% uptime possible? Not in practice. Google's SRE book says 100% is probably never the right reliability target.
Sources
- Google, SRE Book: Availability Table: allowed downtime by availability level
- Google, SRE Book: Embracing Risk: time-based and request-based availability, and the 100% target
- Google, SRE Book: Service Level Objectives: definitions of SLI, SLO and SLA
- Google, SRE Workbook: Implementing SLOs: error budgets and the four-week window
- Google, SRE Book: Monitoring Distributed Systems: black-box monitoring
- AWS, Availability and Beyond: availability, MTBF and MTTR formulas (with the distributed systems and dependencies sections)
- NIST, CSRC Glossary: availability: definitions of availability
- man7.org, uptime(1): the Linux uptime command
- AWS, Amazon EC2 SLA: definition of unavailability and commitments (May 2022)
- Google Cloud, Compute Engine SLA: short outages and credit claims
- Oracle, Cloud Hosting and Delivery Policies: scheduled maintenance excluded from downtime (September 2026)
- AWS, CloudWatch Synthetics canaries: synthetic checks
- Google Cloud, Uptime checks: checks from multiple locations
- AWS, Application Load Balancer health checks: health checks
- Uptime Institute, Tier Classification System and Explaining the Tier Classification System: no availability predictions assigned to Tier levels
- How we run this layer: 24/7 monitoring