AI in practice · 22 September 2026

Why AI pilots fail, and what production takes

Ask any leadership team in the Gulf whether they are "doing AI" and the answer is yes. Ask what is actually running in production, under real permissions, with a number attached, and the room goes quiet. McKinsey's State of AI survey puts regular AI use at nearly nine in ten organisations, and forty per cent of large organisations now report scaling AI agents, up from twenty-seven per cent a year earlier. Yet only thirty-seven per cent attribute any EBIT impact to AI at all, and just six per cent qualify as high performers. Deployment has stopped being the bottleneck. Value has become one.

We spend most of our working week inside that gap, so this note is a field report rather than a think piece.

The model is rarely the problem

When a pilot stalls, the instinct is to blame the technology. In practice the model is usually the most reliable component in the room. The failures live elsewhere:

  • Data that only exists in the demo. Pilots run on a curated extract someone prepared by hand. Production needs that curation automated, governed and refreshed. Gartner has predicted that through 2026 organisations will abandon sixty per cent of AI projects that are not supported by AI-ready data. Our experience says the prediction is generous.

  • No home in the workflow. A system that lives in a separate tab, outside the permissions and approval chains people already work in, gets visited for two weeks and then forgotten. If using the AI takes more effort than the task it replaces, it loses.

  • The verification tax. The moment people feel they must double-check every output, adoption collapses. Trust is built with instrumentation, guardrails and a clear record of what the system did and why, not with reassurance.

  • Nobody owns the distance. Advisers hand over a deck, vendors hand over a licence, and the space between strategy and a running system belongs to no one. That space is where pilots go to die.

What production actually means

Production is not a bigger pilot. It is a different discipline. A production AI system runs inside live workflows, respects the same permissions as the people it works alongside, writes an audit trail a regulator could read, has a measured baseline from before it existed, and has a named person accountable when it misbehaves at three in the morning. Google's DORA research reaches the same conclusion from a different direction: AI acts as an amplifier, magnifying an organisation's existing strengths and weaknesses, and the greatest returns come not from the tools themselves but from the underlying organisational system. That matches what we see. The engineering is necessary; the operational wrapper is decisive.

Scope backwards from live

The practical fix is unglamorous. Choose the smallest system with measurable impact, not the grandest vision with none. Scope the engagement backwards from a system running in live operations inside a quarter. Measure the baseline before you build. Put the dashboard up on day one, not at the steering committee after go-live. A hard deadline forces honesty about all of it.

That is the reasoning behind our ninety-day method: a two-week paid diagnostic to choose the wedge and quantify the case, six weeks to ship the first system into live workflows, and five weeks to harden, govern and hand over. Demos and pilots do not count as done.

If your organisation has more AI pilots than AI systems, the front door is the two-week diagnostic. The whole method is public; there is no fine print behind it.

Related on triway.ai

Begin here

Start with a two-week diagnostic

A paid, two-week working engagement. We map where intelligence compounds in your business and leave you with a plan worth keeping, whoever you choose to build with.

Not ready to book? See where AI would pay off first: twelve questions, three minutes.