The demo worked.

It read the document, found the important information, and produced an answer in seconds. People around the table smiled. Someone said, “Imagine this at scale.”

Then the pilot quietly stopped.

This is common enough that companies have started treating the gap between pilot and production as a mysterious property of AI. I do not think it is mysterious. The pilot and production system are simply answering different questions.

The pilot asks: Can the technology do the impressive part?

Production asks: Can the company live with the whole thing?

The demo gets the clean desk

A pilot usually receives selected data, interested users, patient sponsors, and a team watching it closely. Production receives Monday morning.

Files are missing. Inputs change. A vendor updates a model. The confident answer is wrong. The user ignores the warning. A customer asks why. The person who built the prototype has moved to another project.

None of this makes the pilot dishonest. It makes the pilot incomplete.

Before moving forward, I want direct answers to seven questions:

  1. What exact outcome improved in the pilot?
  2. Compared with what baseline?
  3. Which cases were excluded?
  4. What happens below the confidence threshold?
  5. Who reviews, overrides, and documents the result?
  6. What does one completed case cost at realistic volume?
  7. Who owns the system six months after launch?

If the answers live across five departments, that is already useful information. The project is not only a model. It is an operating agreement.

Accuracy is not the same as usefulness

A team can spend weeks improving a benchmark while avoiding a harder question: is the output useful at the point where someone must act?

Imagine a system that summarises a customer file. Ninety-five per cent of the text may be correct. But the missing five per cent could contain the exception that changes the credit decision, the refund, or the compliance review.

The relevant measure is not only “How often is the summary correct?” It may be:

  • time saved per completed case;
  • errors caught before a decision;
  • uncertain cases routed correctly;
  • customer waiting time;
  • override rate and override reason;
  • cost of review;
  • incidents per thousand cases.

The boring measurement is usually closer to the business.

NIST’s Generative AI Profile explicitly treats measurement, monitoring, incident disclosure, and human-AI configuration as part of risk management. This is helpful because it moves the conversation away from a one-time model score and toward a system that changes over time.

Compliance cannot be the surprise guest

In regulated products, teams sometimes build value first and invite compliance at the end. Then everybody is disappointed when the answer is not a quick yes.

I prefer value and compliance designed together. Not because every experiment needs a legal ceremony, but because the product changes when traceability, data rights, human oversight, or explanation matter.

The EU AI Act’s obligations are arriving in stages. The exact classification and legal duties depend on the use case, and formal advice belongs with qualified counsel. But product teams do not need to wait for a final legal memo to start keeping an inventory, documenting purpose, naming owners, and deciding what evidence they will retain.

Good governance is not a PDF added after launch. It is visible in the workflow.

Production readiness is a product decision

I use a simple readiness check across six areas:

  • Value: a real outcome, baseline, and economic case.
  • Data: lawful access, quality, representativeness, and change monitoring.
  • Failure: known failure modes, thresholds, fallbacks, and stop conditions.
  • People: named owner, trained users, escalation, and human authority.
  • Technology: integration, security, observability, versioning, and vendor dependency.
  • Evidence: logs, evaluation, documentation, and incident handling.

One weak area does not automatically kill the project. It tells us what the next experiment should prove.

Perhaps the next step is not a wider pilot. Perhaps it is a two-week exercise with operations. Or a shadow run where the system makes no live decisions. Or a cost test with ugly data. Or a deliberate pause until the owner exists.

The point of a pilot is not to earn permission to launch. It is to reduce uncertainty.

When the demo works and production says no, do not ask the model to become more impressive. Ask which unanswered question the organisation is protecting you from.

Sources and further reading