
A lot of company software has no way in except the screen. To get data out, something has to click through it the way a person would.

Warehouse safety
warehouse-safety video agent, and a public eval measuring the part everyone skips: whether the verification step earns its place.
sk a vision language model whether footage contains a safety hazard and it will almost always say yes. On my labelled set a single-pass VLM flagged a hazard in 77% of windows that contained none. No amount of prompt rewriting fixes that, because the prompt is not the problem — showing a model a safety camera and asking about hazards hands it an overwhelming prior that hazards are present.
So I treated it as an architecture problem instead. The perception pass is allowed to over-report. Behind it sits a second model that can only remove things: it re-opens the same frames and decides whether the evidence actually meets the bar for the class that was claimed.
The cheap version of this is a reasoning LLM reading the perception pass’s text. That arm scored F1 0.19 — below doing nothing at all. Written evidence is too thin to adjudicate on.
verything runs on NVIDIA-hosted NIM endpoints, so the whole thing reproduces on a laptop with an API key and no GPU: Nemotron Nano VL for perception, a reasoning VLM for verification, and NeMo Retriever embeddings for natural-language search across every analysed window. It also ships an MCP server exposing the timeline as agent tools, so any agent can query the footage in plain language.
Then I hand-labelled 49 windows of real footage and built an eval harness with seven ablation arms — and measured run-to-run variance across repeats rather than quoting a single lucky number.
Three weeks of that project was labelling. There is no shortcut, and an eval built on someone else’s labels would not have caught the failure below.
rame-level verification roughly doubled precision: 0.15 → 0.35 on a single run, 0.34 averaged over three repeats. Recall fell from 0.80 to 0.60, and to 0.50 on the averaged figure. That is a real precision win that costs real recall — not a free lunch, and I would not present it as one.
On a real floor you would tune the verifier’s bar per class. Missing a pedestrian in a forklift path costs more than a missed PPE violation, so they should not share a threshold.
Two findings I did not expect, and would not have got from a demo. Confidence thresholds barely help — raising the bar from 0.15 to 0.80 moved precision 0.15 to 0.18. And text-only verification scores below the naive baseline, which is the result that changed what I built.
Both models described a pedestrian walking in front of a moving forklift. The frame was a black title card reading PASSING IN FRONT OF A FORKLIFT.
either model was hallucinating exactly — they were reading. Burned-in text, slates and lower-thirds are everywhere in real deployed footage, and a system that treats printed words as observed events will fire on all of it. A scene gate in the same call removed the entire class at no extra cost.
This is the argument for hand-labelling rather than sampling. A 0.95-confidence detection on a title card looks identical to a correct one in any aggregate metric.

he eval is small: 49 windows, 10 positive events, one annotator, and that annotator is me. Seven of the ten positives are one class, and blocked_egress has no positive examples at all, so its numbers mean nothing yet. Labels are judgement calls and a second annotator would move them. Treat the precision figures as a directional result on one labelled set, not a benchmark.
A work sample that oversells is worse than one that is small. The limits are in the repo README too, not just here.
The demo was the easy half. The eval is the part that tells you whether the architecture is real, and it is the part that changed what I built.

A lot of company software has no way in except the screen. To get data out, something has to click through it the way a person would.
The rest of it, and what each one cost.

Four New York City databases hold warnings about which buildings are fire risks. None of them talk to each other.