From intersection video to an auditable safety decision.
Caught 3 of 4 AI-labelled near-misses (blind, pre-registered labels); the engineer reviews 12 clips instead of 141 s of video / 51 candidates. Test split: 1 of 2 caught, 6 flags of 25 (split made after v1 saw all events; v3 designed from v1/v2 failures).
CloseCall finds possible near-misses in traffic-camera video for an engineer to review: a tracker measures every close pass, vision models triage the clips for an engineer to review, and the fix memo cites the labelled near-misses. Heads and plates are pixelated in tracked boxes; untracked people may not be. No identity is ever extracted.
Label: near-miss (AI, blind, pre-registered). Model (v3): flagged it · “The car travels straight through the intersection while the pedestrian crosses perpendicularly; neither vehicle nor pedestrian alters their speed or trajectory to avoid the other.”
vote: Qwen3.6-35B-A3B + MiniMax-M3 + gemma-4-31B-it · heads + plates pixelated in tracked boxes
Measured on this run (2026-10-02T19:50:33Z). Footage: CC BY-SA 4.0 public intersection clips from Wikimedia Commons, not a live city feed.
Crash data only arrives after someone is hurt.
A city traffic engineer asking council to redesign an intersection “where nobody has died” has no evidence yet. The camera on the pole has recorded every close call for years, but nobody watches it: near-miss studies mean a person reviewing hours of video by hand, so they are rare and expensive.
Measure every close call. Triage the likely ones.
A cheap first stage (YOLO26 tracking + post-encroachment time) flags every pair of road users that passed through the same spot within 4 seconds. Only those clips go to vision models, which reject most screen-space illusions and flag the likely conflicts, in plain words, for an engineer to review.
Evidence before the funeral, not after.
Near-misses happen far more often than crashes, so they show a dangerous design in weeks instead of years. The ledger turns them into evidence a council can read: the clip, the timing, the explanation, the fix.
Ask in plain words. Get the clips and the memo.
The engineer (or the consultancy doing the study) types “turning cars cutting across pedestrians”, gets the matching clips with timing and explanation, and exports a fix memo that cites every clip. Every step is traced in W&B Weave so the result can be audited.
YOLO26 + ByteTrack
Tracks every car, bus, truck, cyclist and pedestrian; computes post-encroachment time per pair.
NVIDIA Nemotron + embeddings
Nemotron-3 Omni was the video verifier in v1 and v2; nemotron-3-embed-1b powers plain-language search; the memo model is Nemotron-3.5.
VAST Data (adapter, stand-in)
Stand-in on this run: the VastStore adapter (clips, ledger rows, vectors, VSS ingest with a near-miss prompt) is written but not yet run against VAST.
W&B Inference + Weave
Runs the triage vote (Qwen3.6, MiniMax-M3, Gemma-4) on CoreWeave GPUs, drafts the memo, traces every call and scores each verifier against blind labels.
Measured precision of the verifier on this run: 33% versus 9% if every stage-1 flag were reported. See the eval, misses included →