Watch for the Thing You Actually Worry About
Open the manual for a modern "smart" security camera and look at the list of things it can detect. Here is the complete event list from a current 4 MP camera datasheet: motion detection with human and vehicle classification, video tampering, exception, scene change, line crossing, intrusion, region entrance, region exiting. The datasheet is refreshingly blunt about the scope: "the system focuses on human and vehicle targets."
That is the menu. Eight items, most of them geometric. Now think about what actually worries you in your own store.
The menu and the worry list do not overlap
Ask a convenience store owner what keeps them up and you get answers like these. The lottery terminal being used when no customer is at the counter. A beer delivery left at the back door without a signature. Someone eating in the aisle. The back door propped open after ten at night. A line at the register with nobody stepping up to the second till. A cooler door left open overnight.
Not one of those is "line crossing." They are not geometry. They are situations, and a situation is a combination of what, where, when and who-is-not-there.
The gap matters because these are not hypothetical losses. Lottery is a good example. A Colorado grand jury indicted four people over at least 45 scratch-ticket thefts totalling more than $150,000 across sixteen months; the method was to get the clerk away from the counter - a propane exchange, a "stuck" card, spilled gas - and take tickets from the dispenser. In Florida, a compliance sting caught a clerk who scanned a customer's test ticket showing a $1,000 win and kept both the ticket and the claim documentation, selling the claim for $800.
Deliveries are the same story from the supply side. A News4 I-Team investigation into a county beverage delivery operation found crews keeping cases on the truck and filing shortage claims, with two-thirds of businesses shorted, some more than once a week. The fix the county reached for included reviewing loading-dock video - after the fact, which is the only way that has ever been available.
And the internal category is large in aggregate: the National Retail Federation's security survey attributes 29 percent of shrink to internal theft, with an average loss per investigation of $2,180, against total US shrink of $112 billion.
A camera pointed at the lottery dispenser sees all of this happen. The analytics menu just has no word for it.
Why the menu was fixed in the first place
For most of the history of video analytics, a detector had to be trained on a closed list of categories decided in advance. As the paper behind CLIP put it in 2021, "state-of-the-art computer vision systems are trained to predict a fixed set of predetermined object categories". Adding "beer delivery left unsigned" to that list meant collecting a dataset, labelling it and retraining a model - which is why the list stayed at person, vehicle and line.
That constraint is what changed. CLIP was trained on 400 million image-and-text pairs and could classify things it had never been explicitly trained on. Open-set detectors followed: Grounding DINO, published in 2023, detects "arbitrary objects with human inputs such as category names or referring expressions". Vision-language models went further, describing a scene in sentences rather than picking from a list. The practical consequence is that "what to watch for" stopped being a programming task and became a writing task.
Writing the instruction
In our system, you type the thing you are worried about, in your own words, for one camera or for the whole store. "Tell me if the lottery terminal is used and there is no customer at the counter." "Tell me if a delivery is left at the back door and nobody signs for it." "Tell me if the back door is open after ten."
The first AI is already writing down what happens on every camera in plain English. Your instruction is checked against that running description, and when something matches, a second AI pulls the actual clip and confirms it before you are texted. You get the clip and a sentence saying what it saw and why it matched what you asked for.
Two honest caveats. Specific beats vague: "someone is behind the counter who is not wearing a store shirt after closing" works better than "watch for anything suspicious", because the second one is not a description of anything. And the system is not omniscient - a camera that cannot see the terminal cannot tell you about it, and an instruction about something off-camera will never fire. The first thing worth doing is walking the store and asking, camera by camera, what this one can actually see.
Ask for a demo
When getting a new camera-AI service, test this service. Ask the salesperson to type your concern - the real one, in your words - into their system while you watch, and then show you what happens. Not a prepared demo. Yours.
A vendor whose system takes a fixed menu will change the subject to their detection accuracy. That answer tells you what you need to know: the product can watch for their list, and your list is a different list.
Shobdo VideoRAG is an AI agent for the security cameras your store already owns. It writes down what it sees and texts you only when something matters. Learn more or book a conversation.