AI Agents in IT Operations: The Safe Path to Autonomy
Gartner expects 70 percent of enterprises to run IT infrastructure with AI agents by 2029. A practical roadmap for a safe transition to autonomy.

Short answer: AI agents in IT operations are moving from a copilot role to autonomous remediation. Gartner's Predicts 2026 report expects 70 percent of enterprises to operate their IT infrastructure with agentic AI by 2029, up from less than 5 percent in 2025. The same analyst firm also predicts that misconfigured AI will shut down national critical infrastructure in a G20 country by 2028. The right question is not "agent or human" but which decisions get delegated to agents, and behind which guardrails. The answer is a graduated authority model: observe first, then recommend, then approve, and only at the end, bounded autonomy.
From copilot to autonomous operations: the 2026 picture
The numbers show a clear gap between intent and readiness. According to Gartner's Agentic AI Hype Cycle published in April 2026, only 17 percent of organizations have deployed AI agents so far, while more than 60 percent expect to deploy them within the next two years. Gartner describes this as the most aggressive adoption curve among all emerging technologies it measures, and places agentic AI at the Peak of Inflated Expectations. The critical sentence in that analysis: most deployments remain narrowly scoped, and fully autonomous agents are not ready for the majority of enterprise use cases.
At the other end of the gap sits cancellation risk. In a June 2025 prediction, Gartner says more than 40 percent of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. Forrester's 2026 predictions show the bright side of the same coin: in 2026, an agentic AI workflow is expected to autonomously prevent a major system outage using a chain of predictive and causal AI.
These two pictures do not contradict each other; they are two faces of the same fact. Agent technology works, but only in environments where operational discipline already exists. Model quality is not the differentiator; the operating architecture built around the agent is.
The blast radius of automation: the Cloudflare lesson
You do not even need AI to understand the risk profile of autonomous operations; the biggest recent accident of deterministic automation is enough. In the Cloudflare outage of November 18, 2025, a database permission change caused a ClickHouse query to return duplicate rows. The feature file automatically generated by the Bot Management system more than doubled in size as a result, hit a hardcoded 200-feature limit in the proxy software, and the system panicked. The outcome: hours of 5xx errors across a significant part of the internet. The main impact was resolved at 14:30 UTC, and all services returned to normal only at 17:06.
The mechanism here exposes the fundamental risk of every autonomous system: automation propagates wrong decisions at machine speed just as efficiently as right ones. The lessons Cloudflare drew for itself translate directly into the agent era: treat internally generated configuration as untrusted user input, expand kill switch coverage, and review failure modes module by module. If a deterministic pipeline, a system whose behavior is fully specified in advance, needs these guardrails, they cannot be negotiable for a probabilistic AI agent.
Gartner's February 2026 warning points the same way. In analyst Wam Voster's words, the next great infrastructure failure may not be caused by hackers or natural disasters, but by a well-intentioned engineer, a flawed update script, or a misplaced decimal. Gartner's recommendations to CISOs form the skeleton of the guardrail architecture we will build below: safe override modes and kill switches, testing on a digital twin before deployment, real-time monitoring and rollback mechanisms.
The picture in Türkiye: the expertise gap comes before autonomy
In Türkiye, the picture sits a few steps behind the global hype curve, and that does not have to be a disadvantage. According to TurkStat's 2025 ICT Usage in Enterprises survey, 7.5 percent of enterprises in Türkiye use AI technologies, up from 2.7 percent in 2021. The rate climbs quickly with scale: among enterprises with 250 or more employees, usage reaches 24.1 percent.
The really striking data sits on the barriers side. Among enterprises not using AI, 74.2 percent name the lack of relevant in-house expertise as the most important reason, followed by high costs at 67.4 percent and legal uncertainty at 62.4 percent. In other words, the primary obstacle on Türkiye's road to autonomous operations is not the technology, but the human and process layer needed to run it safely. We covered the legal uncertainty dimension, through the lens of the accountability chain, in our article on AI governance; this article focuses on the operations layer.
This picture should be read as a sequencing advantage, not as lateness. Instead of rushing agents into production at the peak of the hype and joining the 40 percent whose projects get canceled in 2027, it is possible to build the guardrail architecture first and open up autonomy gradually.
Guardrail architecture: a four-level authority model
Autonomy is not an on-off switch; it is authority delegated in stages. The model that works in the field has four levels:
- Observe. The agent only reads: it summarizes telemetry, flags anomalies, adds context to incident records. It has no write access to production systems. This level is used to measure the agent's data quality and false positive rate.
- Recommend. The agent produces root cause hypotheses and remediation proposals; a human executes them. Acceptance and accuracy rates of the proposals are measured here. Delegating authority to an agent with a low acceptance rate is deciding on hope instead of measurement.
- Approve. The agent prepares the intervention, and a human approves or rejects it with one click. This is the human-in-the-loop level. The approval screen must show the agent's reasoning, the systems affected and the rollback plan.
- Bounded autonomy. The agent acts on its own within a narrow, predefined area: blast radius constrained, action speed rate-limited, every action paired with automatic rollback, and everything written to an audit log. The kill switch is always reachable and tested through regular drills.
Skipping a level is volunteering to repeat the Cloudflare failure with a probabilistic system. Promotion between levels should be decided by thresholds, not by sentiment: target values for false positive rate, recommendation accuracy and rollback rate should be written up front, and the agent should not be promoted until it sustains those thresholds over a defined period.
Where the agent fits in a managed operations model
Agents can never be better than the operational ground they stand on. Three ground conditions stand out:
- An observability foundation. The agent is only as accurate as the consistency of its metrics, logs and traces. An agent working on unlabeled, scattered telemetry is a fast but unreliable intern. We covered how to build this foundation in our observability guide.
- Machine-readable runbooks. The remediation knowledge in a senior engineer's head cannot be delegated to an agent until it is written down, stepwise and parameterized. A process without a runbook has no automation either.
- An SLA for the agent. The discipline applied to human operators should apply to the agent: targets for decision time, accuracy and rollback rate, with automatic demotion on violation.
These three conditions also explain why this is a managed services operating model question. In a NOC where 24/7 monitoring, incident management and change discipline are already in place, the agent slots gradually into the existing process and gets measured at every level. Without that ground, the agent automates the disorder; and the 74.2 percent expertise gap TurkStat points to turns, at exactly this point, from a technology choice into an operations partner choice.
Five steps for Monday morning
The transition to autonomy starts with an inventory exercise, not a board presentation:
- List your 10 most frequent incident types and mark whether each one has a written runbook.
- Pick the two lowest-risk scenarios that have runbooks and run the agent in observe mode only, read-only, for 30 days.
- At the end of those 30 days, measure the false positive rate and recommendation accuracy; write your promotion thresholds from that data.
- Give no agent write access before the kill switch and rollback mechanism are in place; run the first drill before the agent reaches production.
- Verify that every agent action is written to an audit log that answers who, what, why and when; autonomy that cannot be audited is autonomy that cannot be accepted.
The distance between Gartner's two predictions, the 70 percent operating their infrastructure with agents and the 40 percent whose projects get canceled, hides in whether these five steps were skipped. If autonomy is the goal, safe autonomy is the only viable path.