Conveyors are made for cameras

A conveyor is one of the friendliest places for computer vision: the objects come to the camera, the background is fixed, the lighting can be controlled, and the direction of travel is known. That's why counting, jam detection, and label reading on sorting lines are some of the most reliable vision systems you can build.

It's also where small improvements add up fast. A jam caught a minute earlier or a mislabelled parcel caught before the truck leaves saves real money, every day.

What a demo shows, and what it hides

In a still photo, a model happily boxes every parcel. On a running line, things change: parcels touch and overlap, they move fast enough to blur, they look identical, and the same parcel appears in many frames. Counting boxes in a frame is easy; counting parcels that passed is a different problem.

The demo also hides the question that matters: what does the line do with the result? A count nobody looks at, or a jam alert that goes to an inbox, doesn't change anything.

What a production setup looks like

  • A camera above the belt with fixed lighting and a fast enough shutter that parcels don't blur.
  • A detection model trained on your parcels, then tracking so each parcel is counted once when it crosses a line.
  • Jam logic based on how parcels move: if they stop moving or pile up in a lane, that's an event, not just a crowded frame.
  • Label reading with a dedicated barcode or OCR step on a cropped, sharp image, checked against your sorting system.
  • Integration with the line itself or the operator screen, so an alert actually reaches someone who can act.

Accuracy you can stand behind

Before anyone relies on the counts, I compare them against a ground truth: manual counts for a few shifts, or counts from your existing scanners. That gives an honest error rate per lane, and it shows exactly which situations go wrong so they can be fixed.

This step is boring and it's the one that turns a clever system into one an operations manager trusts.

Where vision-language models fit

On a busy line, a trained detector does the counting; it's fast, cheap, and consistent. A vision-language model is great next to it: describing what's going on in a jam clip for the shift log, or answering "what's blocking lane 3?" in plain words. Each does what it's good at.

Getting started

Pick one lane with a known problem, like jams or mislabels, put one camera on it, and measure. If it works there, more lanes are mostly copy and paste. Try the photo demo with a picture of your own line to see what else it suggests.