The weekend version
Fetch a website, send the text to an LLM, ask for use cases, show the answer. That version took an afternoon and worked on the first few sites I tried. It would also have been a liability on a public page within a day.
Untrusted input, in both directions
Visitors type any address they like. Without checks, a public scanner can be pointed at internal addresses or used to hammer other sites. So it only accepts public domains on standard ports, fetches at most three pages with a timeout and a size limit, and identifies itself honestly.
The page text itself is untrusted too: a website can contain instructions aimed at the model. The prompt treats everything scraped as data, and the answer is forced into a fixed structure, so it can only ever produce use cases, not arbitrary text.
Making it feel fast
A full answer takes around ten seconds, which feels long when nothing happens. The model's answer is streamed, and each finished use case is sent to the page the moment it's complete. You see what the business does after about four seconds and the first use case shortly after, with each real step and a time estimate shown while it works.
For photos, splitting one large request into two that run side by side (objects and ideas) cut the wait roughly in half.
Keeping costs bounded
- Results are cached per website for two weeks, so repeat scans are free and instant.
- Each visitor gets a few fresh scans a day, and the whole site has a daily cap.
- Output length and the model's thinking effort are capped, and photos are shrunk before they're sent.
- Every call logs its tokens, so the real cost per scan is known, not estimated.
Measuring whether anyone cares
A demo that nobody finishes is just a cost. The scanner tracks each step: how many people start a scan, how many get results, how many actually scroll to see them, and what they do next, like emailing the results or trying another site. All of it is anonymous and without cookies.
That's the same funnel I set up for AI features in client products. It's the only honest way to know if a feature is used.
The lesson
The model call is maybe ten percent of the code. The rest is input safety, speed, limits, error handling, and measurement. That ratio is typical, and it's why demos are quick and production takes real engineering.