A demo answers one question

A demo answers "can this work at all?" That is a useful question, and with today's models the answer is often yes within a few days. You point a vision model at a photo of your warehouse, or an LLM at a pile of emails, and it does something impressive.

Production answers a different question: "can we rely on this every day, on inputs nobody picked, when nobody is watching?" That is where most AI projects quietly stop. Not because the model was bad, but because nobody planned for everything around it.

Where the gap actually is

When I look at AI projects that stalled, the model is rarely the problem. The problem is almost always one of these:

  • The data in real life looks different from the demo data: other lighting, other cameras, other email formats, other languages.
  • Nobody decided what happens when the model is wrong, so the first visible mistake kills trust in the whole thing.
  • There is no way to see what it did. When someone asks "why did it flag this?", nobody can answer.
  • Costs were never measured, so the first real month of usage is a surprise.
  • It lives on one person's machine and isn't connected to the systems people actually use.

Decide what a mistake costs

The most important production decision isn't the model. It's what happens when the model is wrong. A missed pallet count is annoying. A missed forklift near-miss is not. A wrongly drafted email that a person reviews is fine. A wrongly sent invoice is not.

Once you know what a mistake costs, the design follows: where a human signs off, which confidence is high enough to act on its own, and what gets logged so you can find the mistakes later. This is also why I like to start with work where mistakes are cheap and get caught by people anyway.

What production-ready actually means

For the systems I build, "production-ready" is a concrete checklist, not a feeling:

  • It runs on real inputs from the real place, and I have tested it on the messy ones, not just the nice ones.
  • Every decision is logged with its input, so any result can be explained and replayed.
  • There are limits: on cost, on how much it can do without approval, and on what it can touch.
  • Someone gets told when it breaks, and it fails in a safe direction.
  • There is a simple way to measure if it is getting better or worse over time.

The demos on this site are honest about this

The website scanner and the photo demo on this site are real, running on Gemini. They are also deliberately demos: they show what is possible for your business in seconds. Getting from those ideas to something that runs on your cameras or inside your systems is exactly the part I do.

The articles in this series take each example from those demos, like forklift safety, parcel sorting, the loading dock, robot cells, and AI agents on top of existing software, and walk through what it really takes to put them into production.

How I approach it

I start small and real: one camera, one workflow, one team. I get it working on real data quickly, put a human in the loop, measure what it gets wrong, and only then widen it. It's slower than a demo and much faster than a project that never ships.

If you have a demo that never made it, or an idea you want to test properly, that's a good conversation to have.