All notes

AI in production

Why enterprise AI stalls, and what we've seen work

A tangled mass of thread, with one green strand pulled out and threaded cleanly through a row of steel guides.

Ask ten people whether AI is delivering in large enterprises and you'll get ten answers. McKinsey's State of AI in 2026 survey (August 2026) found nearly nine in ten organisations now use AI regularly. Only 37% see any impact on profit, and just 6% get significant value from it. So almost everyone is using it, and very few are getting much out of it.

Meanwhile the companies that got it right are pulling away. From where we sit, building AI into banks, insurers and healthcare businesses, it mostly comes down to how you go about it.

Why so many rollouts stall

  • The problem isn't clear. AI gets pointed at whatever is most visible, not at what it's actually good at. Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027, with unclear business value one of the main reasons. Some problems need a spreadsheet, not an agent.
  • "Use AI" becomes a KPI. One survey of nearly 1,300 companies found 58% require some employees to use AI. As one security chief put it, the question isn't "use more AI", it's "use more AI to do what?" Without an answer you get lots of activity and very little change to how work gets done.
  • Proof of concept to production is a long road. A demo on clean sample data takes a week. Production means real data, security reviews, old systems, monitoring and support. That's where most projects die. Gartner's view is blunt: most agentic projects today are experiments and proofs of concept that will struggle to scale beyond the prototype.
  • Regulated industries have extra hurdles. In the Bank of England and FCA's survey, firms ranked data privacy as the biggest AI risk and data protection as the biggest regulatory constraint. Nearly half said they only partly understand the AI they use. If a model can't show where a number came from, it won't survive an audit.
  • Cost. If every task goes through a frontier model, the bill is hard to justify. Most work doesn't need that. The Stanford AI Index found the cost of GPT-3.5-level performance fell 280-fold in under two years. Open-weight models are now within a couple of percentage points of closed ones. Agents use far more tokens than chat, though, so routing work to the right model matters more, not less. We wrote about how we handle this in Some things we've learned scaling LLM inference.
  • The data isn't ready. More on this below, because it's the big one.

What we've seen work

1. Let the business lead, with technology alongside

The best projects are owned by the people who do the work and feel the pain. Technology makes it secure, supportable and real. BCG's rule of thumb is that about 10% of AI's value comes from the technology, 20% from data and algorithms, and 70% from changing people, processes and ways of working.

2. Pick a problem small enough to finish

It still has to matter, but it should be small enough to reach production in weeks, not quarters. AI is moving faster than anything we've worked with, and short cycles let you learn before the ground shifts again.

3. Build for change

Six months used to be a short cycle in enterprise technology. In AI it's a lifetime. In September 2026 alone, OpenAI released GPT-6 Astra and Anthropic released Claude Fable 5.1. Three weeks later Anthropic followed with Opus 5.5, which beat Fable 5.1 on several agentic benchmarks at 60% less. Google closed the month with Gemini 4, which it claims beats all of them. That's four frontier releases in four weeks, and each one shook up the rankings. Whatever model you choose today will probably be overtaken before your project goes live. That has never been true of enterprise technology, and most organisations are still getting used to it. Keep models swappable, keep your prompts and evals in your own hands, and don't lock a workflow to one provider.

4. Get the data foundation right

This matters more than ever. An agent is only as good as what it can see. Gartner found that 63% of organisations don't have, or aren't sure they have, the right data practices for AI. It expects 60% of AI projects without AI-ready data to be abandoned through 2026.

In practice that means:

  • Clean, standardised data taken from source, with one agreed golden source for each thing that matters.
  • Controls and reconciliations built in, so bad data fails loudly instead of quietly becoming the baseline.
  • An ontology or semantic layer that describes what the data means in business terms, so a model, or a new analyst, can reason about it.
  • Clear ownership of definitions, lineage you can trace, and access rules that respect privacy.

It's the least exciting part of the work, and it decides whether everything after it works. Get it right and every new use case gets cheaper. Skip it and every new use case starts from scratch.

5. Start by talking to your data

You don't need to begin with a fleet of agents. People are used to asking questions. Nobody wants to scroll through a 60-page PDF or click through five dashboard tabs to find one number. An assistant that answers questions over your own data is useful from day one. It also tests your data foundation fast, because without a proper semantic layer it will give confident wrong answers.

6. Map the work before you automate it

The real value comes from role-specific agentic automation: agents that do a defined part of a defined job. To build those, you need to understand the function in detail. Who does what, in which systems, which decisions they make and where things go wrong. McKinsey found nearly three-quarters of its AI high performers had fundamentally redesigned their workflows, up from 55% a year earlier. That's the single clearest pattern in their data.

7. Don't start with "replace the team"

There's a lot of nuance in what people do, and much of it isn't written down anywhere. Klarna famously moved most of its customer service to AI, then started hiring people again when quality slipped. In McKinsey's data, 32% of companies expected AI to cut their headcount, but only 14% actually saw it. Agents create value by letting a team handle far more volume, giving clients a better experience, or taking the grind out of internal support. For now, someone still has to manage the agents, check their work and handle the exceptions. Keeping a person in the loop is part of the design, not a stopgap.

8. Accept that some of it is throwaway

Building is getting cheaper every month. What used to take months to prototype now takes weeks. So try things, keep what works and bin what doesn't. Some of what you build this year will be replaced next year by something better, and that's fine.

Where this leaves you

If you're not yet embedding AI into everyday work, you're not alone, but you are falling behind. According to McKinsey's 2026 survey, the share of companies scaling AI across the business rose from 38% to 44% in a single year. Among companies with more than $1 billion in revenue, the share scaling AI agents jumped from 27% to 40%. They're working out how to add AI capacity to their people instead of chasing headcount cuts, and the ones that do it well are seeing real returns. We see it first-hand in our work, so we know it's real. The ones getting ahead didn't start with the biggest plan. They started with good data, a real problem and a willingness to keep iterating.

Working on something like this?

If any of the above sounds familiar, we have probably run into it before.

Start a conversation