CIO Report Series: Why a vast majority of agentic AI investments still fail

CIO Report Series: Why a vast majority of agentic AI investments still fail

CIO Report Series: Why a vast majority of agentic AI investments still fail

Kelly Rupp

Kelly Rupp

Kelly Rupp

Product Marketing

Product Marketing

Product Marketing

Share

CIO Report 2026 series: insights from exclusive report, based on interviews with IT leaders at Cloudflare, Cursor, Lyft, GitLab, Nextdoor, Zip, Rubrik, IMC Logistics, TrendAI, Palo Alto Networks, and Zscaler.

According to Gartner, 79% percent of AI technology purchase decisions end in regret. Most organizations deploy what they buy. The regret comes after, when the investment turns into a real tool, producing real results… which don’t live up to what the demo showed.

Agentic AI makes this worse. An automation that produces bad outputs fails on a single task. An agent that makes wrong decisions fails loudly across every ticket, every workflow, and every interaction it touches.

Regret. Rinse. Repeat.

Manu Narayan, CIO at GitLab, has watched IT organizations fail from both ends of the spectrum. The first is the extended evaluation cycle. "If you take a six- to twelve-month RFP cycle to make a decision on a platform," Narayan says, "you're going to make a decision on data that's 12 months old."

The second failure is faster, but more expensive. Organizations that buy a platform quickly and deploy it without defining the problem it's solving. Capability purchased. Foundation skipped.

Agentic AI makes decisions on behalf of people and organizations. Deploying a decision-making system on top of an unclear process produces inconsistency at the speed of automation.

A demo environment never surfaces this. Sandbox deployments run against curated data, optimal conditions, and cases the vendor chose to demonstrate. Production deployments run against the full complexity of real operations: ambiguous tickets, missing fields, edge cases that don't appear in any training set.

What the failure actually looks like

The gap between demo and production is where most agentic AI investments collapse.

Rachel Jin, Chief Platform and Business Officer at TrendAI, has watched this play out across her own organization as it built what she calls Trend IQ, an AI-driven platform running on all of TrendAI's internal data. Her read on why teams struggle: "I feel they are all scared. Scared if they are not using AI. But they are also a little bit worried how they can use AI." Urgency without clarity is a recipe for failure. 

Ultimately, that means taking the time to build a strong foundation will determine if your agentic deployments succeed or fail. And that happens long before the agent even runs.

IT teams that have run successful agentic deployments share a common prior step: they cleaned up the process before they automated it. Not perfectly. But well enough that the escalation paths were clear, the categories were defined, and the cases the agent couldn't handle had a known destination.

Joel Tracy runs IT at IMC Logistics, a mid-size operation where the team is small and the operational complexity is high. His view of what AI is actually for is precise: "The idea is you want to automate the things that are repetition based and repetitive so that your employees can spend more time managing the exceptions, not the norm." That clarity determines the scope. His team's successful deployments started narrow, with one defined category and one clear metric, then expanded. The ones that didn't work skipped the scoping step, deployed broad, and produced no legible signal about what to do next.

What separates the deployments that work

The organizations measuring real results made one decision differently. They defined what the agent was replacing before they deployed it.

Piru Chheang, Head of IT at Zip, reached 50 percent ticket deflection in production with a team of one and a half serving over 800 people. He described his starting point plainly: "We're more reactive than proactive at this point." His team measured deflection rate against real inbound volume from day one. The success criteria were set before the deployment ran.

This is one of the central patterns in the IT's Time to Build report: the IT leaders measuring real outcomes from AI started with a clear problem definition and a legible metric, not a capability purchase.

The measure that matters in agentic AI is resolution rate in production against real inbound volume. Not demo accuracy. Not vendor-provided benchmarks. That metric tells you whether the agent's decisions are correct. Everything else is theoretical.

Eighty-two percent of IT professionals have realized tangible value from AI investments, according to the State of AI in IT 2026. Only 20 percent have AI fully embedded across service management teams. The gap between those two numbers reflects deployment depth. Buying is easy. Getting to production at real scale requires the foundation work most organizations skip.

The question before every agentic deployment

The practical test for any agentic AI investment is a single question. Is the process being automated actually understood well enough to automate?

If yes, escalation paths are clear, edge cases have owners, success criteria are set, and the deployment will produce results. If you aren't sure, spend your first dollar clarifying the process, not buying the platform.

For a closer look at how to scope and run a first deployment that generates real data, see How to Run Your First AI Pilot in 30 Days.

Narayan's observation about the fear that slows AI adoption extends here: "There's a lot of fear of the unknown." Teams that skip foundation work often do so because the foundation work is harder to sell internally than the capability. An agent platform has a demo. Process clarity does not.

The organizations getting results deployed the same tools in the same market. They started with a cleaner answer to what the agent was supposed to replace. Sequencing was the variable.

For more on what IT teams do with the hours that successful deployments return, see Escaping the Reactive Trap.

The IT's Time to Build report goes deeper on every layer of this framework, with insights from IT leaders at Lyft, Cloudflare, GitLab, Zip, and more. [Download the report →]


Frequently asked questions

Why do most agentic AI investments fail to produce results?

Most deployments skip foundation work: defining the process the agent will run, setting escalation paths, and establishing a clear success metric before deployment. Agents making decisions on top of unclear processes produce inconsistency at scale.

What should IT teams do before deploying an agentic AI system?

Define what the agent is replacing. Set success criteria before deployment. Identify the escalation path for cases the agent can't handle. The scoping work takes days; skipping it costs months.

How do you measure whether an agentic AI deployment is working?

Resolution rate in production against real inbound volume. Not demo accuracy, not sandbox performance. The metric that matters is how many real cases the agent resolves correctly without human escalation.

What makes agentic AI different from standard automation?

Agentic AI systems own decisions, while standard automation executes rules. Failure modes are harder to predict and wider in impact: an agent making wrong decisions fails across every case it touches, at whatever speed it runs.

Is 98% really the failure rate for agentic AI?

Gartner's figure, 79% of AI technology purchase decisions end in regret, covers AI broadly. Agentic AI specifically carries higher risk when deployed without clear process foundations. The number reflects purchase regret, not deployment failure, but the pattern is consistent: capability without foundation produces poor outcomes.

Subscribe to the Console Blog

Get notified about new features, customer
updates, and more.

Your IT team could run like this too

Your IT team could run like this too