What the photo demo on this site does
When you drop a photo into the demo on this site, it goes to Gemini, a vision-language model. Two requests run side by side: one finds the important objects and their boxes, the other describes the scene and suggests use cases. It works on almost any photo, with no training, which is exactly why it's the right tool for a demo.
It's also a good illustration of the limits. The boxes are decent, not precise. It takes a few seconds per photo. And it costs a fraction of a cent per image, which is nothing for a demo and a lot for 30 frames a second from twenty cameras.
What trained detection models are good at
A detection model like YOLO, trained on images from your own site, does one narrow job: find these specific things, very fast, very consistently. It runs in milliseconds on a small GPU, works offline next to the camera, and gives the same answer for the same input every time.
That makes it the backbone of anything that runs continuously: counting parcels, tracking forklifts, checking every product on a line. The price is the labelled data and the training work up front.
What VLMs are good at
Vision-language models shine where the question is open or changes often: "is anything blocking this aisle?", "what went wrong in this clip?", "does this shelf match the planogram?". They need no training data, understand context, and answer in plain language.
They are great for prototypes, for low-volume checks, and for explaining what a detector already found.
How I usually combine them
- Start with a VLM to prove the idea on real photos in days, before investing in data.
- Use the VLM's output to speed up labelling a dataset for a fast detector.
- Run the trained detector on every frame for counting, tracking, and alerts.
- Send only interesting moments (events, anomalies, stops) to the VLM for a description or a second opinion.
- Measure both against a hand-checked set, so you know where each one fails.
Cost and speed in production
For continuous video, a trained model on an edge device is almost always cheaper and faster, and the video never leaves the building. VLM calls are best kept for the moments that matter. In the demo on this site, every request has hard limits: the photo is shrunk before upload, output is capped, and usage per visitor and per day is limited and tracked. The same discipline applies in any real system.
The short version
Use a VLM to find out what's possible and to explain things. Use a trained detector for anything that runs all day. Most production systems I build end up with both.