Everyone Runs Prometheus: So Why Do We Still Hear About Incidents Late?
The observability tooling debate is over; the problem has moved. A measurable framework for noise, cost and alert fatigue.
Short answer: The question of which tool to install is settled. In CNCF's 2025 annual survey Prometheus is used by 77% of respondents, making it the most widely adopted CNCF project after Kubernetes. Teams still learn about incidents late, though, because the bottleneck is not the tool: according to Grafana Labs' fourth annual observability survey, the single biggest obstacle to faster incident response is alert fatigue, named by 30% of respondents, nearly double the next answer.
This article deals with the work that starts after the stack is running: deciding what becomes an alert, who it reaches, and what that costs.
Why the tooling debate is over
The numbers all point one way. In CNCF's survey, Kubernetes use in production reached 82%, and Prometheus became the de facto standard alongside it. Grafana's third annual survey shows the same picture from another angle: more than two-thirds of organizations run Prometheus in production, with a further 19% evaluating it or building a proof of concept. 70% of respondents use Prometheus and OpenTelemetry together.
So "which metrics engine" is no longer a choice but an assumption. An organization setting up observability today takes Prometheus as a given and debates what to put on top of it.
Where the problem moved
The same survey draws on 1,363 responses from engineers, reliability engineers and technology leaders across 76 countries. Their concerns for 2026 have nothing to do with tool selection:
- Complexity and overhead: 38%, top of the listSignal-to-noise: 34%Cost: 31%
All three have the same root cause. When collecting metrics is cheap and easy, teams collect everything; every series collected produces a storage cost, a query cost and an attention cost. Dashboards fill up but nobody reads them, alerts arrive but nobody opens them.
The most practical finding in the survey: the number of observability technologies organizations use fell from nine to eight. Mature teams are removing tools rather than adding them.
Alert fatigue is a design problem, not a culture problem
"The team ignores alerts" is almost always the wrong diagnosis. If an engineer has found the 3am page unnecessary three times, assuming the fourth is unnecessary too is rational behavior. What needs fixing is the alert itself.
Three rules make a difference, and all three are measurable:
Every alert needs an owner. An alert with no owner is the one everybody sees and nobody opens. We covered how to write ownership down, step by step, in our incident escalation matrix piece.
Every alert must map to an action. "CPU exceeded 80%" produces no action. "The queue has not drained for five minutes and customer requests are backing up" does. The second measures an outcome, not a symptom.
Thresholds must derive from your own targets. A vendor default does not know your workload. As with recovery targets, the right direction is bottom-up: calculating RPO and RTO from business cost also tells you which delay genuinely deserves a page.
Why cost is first, not third
In the same survey, cost is the single most important tool selection criterion for the third year running (65%). That marks observability's move from an engineering topic to a budget topic.
Cost is usually driven less by data volume than by cardinality: every label added to a metric multiplies the number of time series stored. A label with unbounded values, such as a user ID or a request ID, can split one metric into millions of series on its own. That inflates the bill and the query time at the same time.
The approach that works in practice is to treat retention as tiered rather than as a single number: high-resolution data for a short window, summarized data for the long one. What to keep and for how long is a cost decision, and for regulated records you also need to check the statutory retention periods.
Where to start
Answer three questions before installing another tool:
- Of the alerts fired in the last thirty days, how many led to an action? If fewer than 20% did, the problem is your thresholds, not your visibility.What are your five most expensive metrics, and into how many series do they split? The answer is usually a short list, and the fix is often just removing a couple of labels.When an alert fires, is it written down who opens it? If not, the alert is just a notification with no process behind it.
Answering those three changes more than any dashboard will. You can see how Eclit builds this layer on the observability and APM page.
The point of observability is to get a problem to the right person before the customer notices, not to collect more data. Prometheus makes that possible, but it does not guarantee it.