Video In, Text Out: What Changes When Footage Becomes Words - Shobdo Blog

Video In, Text Out: What Changes When Footage Becomes Words

· Shobdo Team
Video In, Text Out: What Changes When Footage Becomes Words

A security camera is a machine for producing more footage than any human will ever watch. Point four of them at a store and every day they record ninety-six hours of video, all night included. The store is probably open only for eleven hours. The useful part of that day might be forty seconds long, and you will not know where those forty seconds are until something makes you go looking.

The interesting question is not how to store all of it. It is what to keep instead.

Nobody is watching, and that was always true

The idea that cameras are watched is a comfortable fiction. Work done at Sandia National Laboratories for the Department of Energy, and published by the National Institute of Justice, found that "after only 20 minutes of watching and evaluating monitor screens, the attention of most individuals has degenerated to well below acceptable levels". That finding is old enough to be quoted back by the Government Accountability Office in congressional testimony, and nothing since has repealed it. A survey of security integrators concluded that less than one percent of all cameras are ever really monitored live.

The storage side is equally lopsided. One 30 frames-per-second stream fills a 6 TB drive in about 84 days with H.264 encoding, and a modest multi-camera recorder easily needs more than 10 TB. That is a lot of disk devoted to empty aisles, held in case someone eventually needs ninety seconds of it.

What a written record does that video cannot

Now imagine the same day as text: a timestamped list of what happened, written in plain English as it happened. Someone came to the register at 2:14. A delivery arrived at the back door at 6:40 and left at 6:51. A person in a grey jacket walked the length of the frozen aisle three times between 4:05 and 4:12.

A 24-hour activity timeline for one store camera above a list of timestamped observations, each a plain-English sentence such as 'The person reaches the foreground near the fish counter, still drinking from the white object.'
One camera, one day, in our dashboard. Every bar is a stretch of activity; behind each one is a time and a sentence.

Four things change immediately.

It is searchable. Finding "the delivery on Tuesday afternoon" in text takes a second. Finding it in video takes a person, a scrub bar and a bad hour. The research community has been working on this translation for years - the task of detecting and describing every event in a video in natural language was formalised as "dense-captioning events in videos" in 2017, and models that output timestamps and descriptions in a single sequence followed. What was a research benchmark is now something you can run over a store's cameras.

It fits in a text message. A sentence describing what happened, with a short clip attached, can reach an owner's phone while he is at his other store. Ninety minutes of video cannot.

It is small enough to keep. Text costs a rounding error to store, which means a store's history does not have to be sacrificed to disk space. The video stays on the recorder, where it always was.

It leaves the video where it belongs. This is the part we care most about. If what travels is text, then video is not stored in the cloud, nobody at the vendor has a library of your store to browse, and the privacy question stops being a promise and starts being a property of the design.

The clip is the evidence. The words find it.

When something goes wrong, the people who ask about it want proof, and proof means video. A sentence in a log is not evidence; anyone can type one. But someone still has to find the right clip, and someone still has to write down what happened for the police report. The written record does both of those jobs.

The Insurance Information Institute's guidance on filing a business claim is explicit about paperwork: for a crime loss, "contact the police and obtain a copy of the police report", prepare an inventory of what was damaged or taken, and expect a signed, sworn proof of loss. A prosecutor's primer published by the Department of Justice's Bureau of Justice Assistance notes that "some estimate that video evidence is involved in about 80 percent of crimes" and that charging decisions often have to be made in a 24-to-48-hour window.

In every one of those situations, what closes the gap is a written account of what happened and when, with the clip that shows it. A store that already has the first half has done most of the work before anyone asks.

What this looks like in a store

In a 40-camera grocery store in Massachusetts running our system today, the day's record is a page you can read. Each entry says what the camera saw and, when something crosses a line the owner cares about, the owner gets a text with a clip and the reason it was sent. Nothing else leaves the building. The recorder keeps the video as it always did; if a clip is deleted there, it is gone, because there is no second copy.


Shobdo VideoRAG is an AI agent for the security cameras your store already owns. It writes down what it sees and texts you only when something matters. You can say in plain English what you want it to monitor and send alert about. Learn more or book a conversation.

Surveillance AI VideoRAG Privacy Retail