All insights

AgentCore in the wild: where agents pay back, and what separates the 31% from the rest

TL;DR — AI agents are delivering measurable value in real businesses today, but only where they're carefully placed: constrained workflows, a metric you already track, a clear owner. Around 40% of agentic projects are expected to be cancelled by 2027, and the failures are placement, engineering and identity — not model capability. The production successes all share the same three disciplines, and each maps to a decision you can make before spending anything on platforms. Andy Rea's hands-on AgentCore series covers the how — five posts so far, with more on the way; this piece covers whether and where.

If you're weighing whether AI agents deserve a place in your business, the honest answer is: probably yes — but almost certainly not where the hype is pointing. Agents matter because they're the first form of AI that does work rather than assists it: software that runs a whole workflow against your systems and moves a number you already report. The deployments in this piece show what that looks like when it's real — support tickets down 75%, campaign builds 30% faster, value proven in months — and every one of them earned it the same way.

Andy Rea is working through every major AgentCore service hands-on and writing up exactly what each one gives you, and what it doesn't — five posts so far covering runtime, tools, memory, code execution and identity, with more of the platform still to come. What follows sits alongside that series and tackles the buyer's side of the question: where would an agent actually pay back for you, and what do the businesses getting real value do differently?

Most agents never reach production. Yours can.

The numbers set the stakes. Roughly 80% of enterprise applications now embed an agent in some form; only about 31% of organisations run one in production.

Embedding is easy. Operating is hard.

And Gartner expects over 40% of agentic projects to be cancelled by 2027.

The drop

Share embedding an agent, against the share operating one in production

Embed an agent in some form Roughly 80% of enterprise applications embed an agent in some form ~80% Operate one in production Only about 31% of organisations run an agent in production ~31% The drop: what the post-mortems blame placed on the wrong workflow engineered as a demo identity deferred until audit
Roughly 80% of enterprise applications embed an agent; about 31% of organisations operate one in production. Gartner expects over 40% of agentic projects to be cancelled by 2027. Sources are listed at the foot of the article.

Here's what should give you confidence rather than pause: the post-mortems don't blame the technology. They name poor integration, unclear ownership, and production concerns — identity, permissions, auditability — deferred until the pilot hit a wall. Every one of those is a decision, not a limitation. Which means the 40% is avoidable, and avoiding it costs judgement, not budget.

Pick the workflow before you pick the platform

The businesses getting real value aren't the ones running the most pilots.

They run two or three high-value, production-shaped use cases with clear business owners, defined KPIs and explicit guardrails — and they choose workflows where the metric already exists, so the before-and-after is undeniable. Agents are going mainstream in constrained, well-governed domains — IT operations, employee service, finance operations — not in open-ended transformation programmes.

Start where mistakes are cheap

IBM's AskHR — built on IBM's own platform, not AgentCore — resolves 94% of common employee questions on its own and has cut support tickets by 75%. The internal helpdesk is a strong first agent because the failure mode is benign: the data is owned, and a wrong answer costs a minute, not a customer.

Weigh speed-to-learning

When you weigh a placement, weigh how soon you'll know. As an example, agents tied to a revenue number the business already watches — sales development — show whether they're working in around 3.4 months; finance and operations agents average nearer nine. That's no verdict on those placements — it's a reminder that time-to-value differs by workflow. A fast proof cycle tells you sooner whether a placement was right, in either direction, which is why speed-to-learning is one of the factors we consider when assessing placements.

Time to prove value, by metric

Months for an agent to return its cost, by the metric it's tied to

Sales-development agents Sales-development agents: cost returned in around 3.4 months ~3.4 months Finance and operations agents Finance and operations agents: cost returned in nearly nine months ~9 months 0 3 6 9 months
Agents on revenue-proximate metrics evidence value fastest. Slower payback isn't failure — but a visible metric shortens the proof cycle. Sources are listed at the foot of the article.

Apply the buyer's test

For a buyer, that's the useful test. If you can name the workflow, the number it moves, and the person accountable for that number, an agent is worth assessing. If you can't yet, the project isn't ready to start — and that's far cheaper to learn now than after the build. The placement work comes first, it's a business judgement you're already equipped to make, and it's exactly what our Placement Assessment exists to pressure-test.

Do what the 31% do

The named production stories are more instructive than any launch-day customer list, and they span sectors. Not all of them run on AgentCore, either — the disciplines are what these deployments share, whatever the platform.

  • Start where the metric is visible. Epsilon, part of Publicis Groupe, expects to cut campaign build times by up to 30% with automated campaign design, targeting and optimisation on AgentCore — build time was already a number they managed.
  • Layer over what you own. Innovaccer turned its existing healthcare APIs into agent-compatible tools through AgentCore Gateway with compliance intact; KTern.AI did the same over SAP. Neither ripped anything out — a pattern that matters if you have a heavy existing estate.
  • Keep the workflow constrained and the data owned. Amazon's Devices supply chain agents work from product specs the team controls — generating quality-control test procedures and training vision systems on the manufacturing line, not open-ended autonomy.
  • Regulated doesn't rule you out. Itaú Unibanco, one of Latin America's largest banks, is building secure, personalised digital banking on AgentCore — this approach survives bank-grade scrutiny.
  • Scale by multiplying small agents, not building one giant. JPMorgan runs more than 450 agentic use cases in production on its own internal platform — many small, governed applications rather than one do-everything assistant. When agents work, this is how they grow.

Learn from Cox Automotive

The most useful case is Cox Automotive, because they published lessons rather than a press release: log everything, iterate small, and accept that core software engineering still applies — with proprietary, well-governed data as the real differentiator, not the agent framework. Nothing on that list is exotic. That's the point: this is deliverable with disciplines your organisation may already have.

The three disciplines of the 31%

What the production successes share, and the Levantar pillar each one maps to

01 Placement a metric you already track a named owner AI VALUE DIAGNOSTICS 02 Engineering logged iterated governed data AGENTIC AI 03 Identity know who the agent acts for AI & LLM SECURITY
Placement, engineering and identity. Each is a decision you can make before spending anything on platforms, and each maps to one of the three things we do.

Price in the platform trade-offs

Frameworks stay portable at the top, but the operational layer underneath — sessions, memory, identity, observability — belongs to the platform, and that's where lock-in forms. Here that platform is AWS, but the trade-off comes with any managed agent runtime — the point is to price it in rather than discover it. And AgentCore is deliberately unopinionated about reasoning: it runs whatever agent you give it, well-designed or not. It removes infrastructure work, not engineering judgement — which is precisely why the disciplines above decide the outcome.

Don't let confidence outrun your controls

One finding should shape every buying conversation. In Gravitee's 2026 research, confidence in agent visibility rose from 82.6% to 91.8% — while the proportion of agents actually being monitored and secured moved only from around 47% to 52% as estates roughly doubled. And only one in five say all of their agents are fully secured and governed before going live — roughly eight in ten organisations are shipping agents to production without full pre-deployment controls. Meanwhile 54% of organisations report a confirmed or suspected agent security incident in the past twelve months.

Confidence vs coverage

Gravitee 2026: what respondents believe about agent visibility, against what is actually monitored

2025 2026 82.6% Confidence in agent visibility, 2025: 82.6% 91.8% Confidence in agent visibility, 2026: 91.8% ~47% Agents monitored and secured, 2025: around 47% ~52% Agents monitored and secured, 2026: around 52% 2025 2026 2025 2026 Confidence in agent visibility Agents actually monitored and secured Agents fully secured and governed before go-live, 2026 19.7% 80.3% All agents fully secured and governed before go-live: 19.7% Ship some agents without full controls: 80.3% all agents fully secured and governed ship some agents without full controls
The estate is growing faster than the controls. Only 19.7% say all agents are fully secured and governed before going live.

Solve identity first

The root issue is identity. An agent isn't a service account: it selects tools contextually, acts on behalf of different users, and can delegate to other agents.

Machine identities already outnumber human ones by 45:1 to 100:1, and 24 million leaked machine credentials were found on GitHub in 2025 — 70% of the 2022 ones still valid.

The risks concentrate in four known places — over-broad scopes, token theft, prompt injection coercing an agent's legitimate permissions, and telling agent actions from user actions — and each has a known mitigation. The usual failure is skipping the mitigation, not ignorance of it.

Ask this one question

For a buyer the question is simple and worth asking any vendor or internal team: who is this agent acting for, and can you prove it after the fact? If the answer is vague, you've found your risk. It's the question Andy Rea's fifth post answers hands-on, and the ground our AI & LLM Security work covers.

Start with the hands-on series

If the placement case stacks up, the series is the practical route in — each post deployed, invoked and poked at, including the parts that didn't behave. Five so far:

  1. What AgentCore actually is, and getting a first agent running on it
  2. Giving your agent tools with AgentCore Gateway
  3. Giving your agent memory that survives the session
  4. Letting an agent run code, without letting it run loose
  5. Knowing who your agent is acting for, with AgentCore Identity

More are on the way as Andy Rea works through the rest of the platform — the series will run to around eleven posts covering AgentCore end to end.

Get into the 31%

Agents can work for you. The evidence says they work when they're placed carefully and meaningfully — on a workflow you can name, with a number you already track, engineered like real software, with identity solved before go-live. Get those right and you're in the 31%.

If you want to test whether a workflow of yours passes that bar, the first conversation is diagnostic, not a pitch.

Sources

Talk to us about this

Contact

Talk to us before your next AI investment.

We run this work in our own business every day. If you are planning to run it in yours, the first conversation is diagnostic, not a pitch.