AI strategy

Beyond the pilot: what it takes to put AI into production

The demo is the easy part. What separates a pilot from something the business depends on is mostly unglamorous.

Updated September 2026

The short version
  • Pilots stall because a demo has no permissions model, no evaluation, no monitoring and no owner — and production needs all four.
  • An assistant inherits the permissions of whoever asks. That makes the content estate part of the security design.
  • Without a way to measure whether answers are right, you cannot tell a good release from a bad one.
  • Someone has to own it after go-live, or it degrades quietly while everyone assumes it is fine.

Almost every organization now has an AI pilot that works. Far fewer have one people depend on. The gap is not model quality. It is everything a demo does not have to do.

Why pilots stall

A pilot is built for one audience, on content someone curated, with the builder in the room when it answers. Production has none of those things. It answers people with different permissions, on content nobody curated, when the builder is on holiday. Each of the sections below is something the pilot got away with skipping.

Identity and permissions

An assistant grounded in company content will answer according to what the person asking is allowed to see — which means your permissions model is now a security control on AI output. Where permissions have drifted, the assistant does not create the exposure, it reveals it.

Before production: know what “everyone” can currently reach, close the obvious oversharing, and decide which content is in scope at all. This is usually the longest item and it cannot be compressed by adding people.

Grounding, and what happens when the source is wrong

The quality of an answer is mostly the quality of what it read. Three versions of a policy, one current, none labelled, produces confident answers from the wrong one.

Production needs a defined source set, an owner for each part of it, and a rule for what happens when content goes stale. Citations on every answer are worth insisting on — they are what makes a wrong answer diagnosable instead of mysterious.

Evaluation

You need a way to tell whether a change made things better. That means a set of real questions with known-good answers, run before each release, with the results recorded.

Without it, every change is a guess and every complaint is an argument. This is the single most common thing missing when a pilot has been running for months.

Monitoring and cost

Two questions have to be answerable at any moment: is it working, and what is it costing. Usage-based billing means a successful rollout and a budget problem look identical in the first week.

Set a ceiling and an alert before go-live, not after the first invoice.

Support and ownership

Someone owns it. Not the project — the service. They handle the “it gave me a wrong answer” tickets, they decide when content needs fixing, they watch the evaluation results, they hold the cost.

An assistant with no owner degrades quietly. The content ages, the answers get worse, people stop using it, and eighteen months later the licences are still being paid.

Change, and the part people skip

The last piece is the least technical. People do not adopt a tool because it exists; they adopt it when it beats what they do now on a task they actually have. Rollout should be to real teams on real work, with usage measured rather than assumed, and with the first group chosen because their work suits it rather than because they volunteered.

What this means for a business case

If a pilot cost very little, the production version will not. Most of the cost sits in content remediation, evaluation, and the person who owns it afterwards. A business case that prices only the build will be wrong by a wide margin, and the overrun will arrive in the quarter after go-live.

Stuck between a working demo and a live service?

Bring the pilot. We will tell you what stands between it and production, and what that is worth doing.