A flight recorder for intersections

From intersection video to an auditable safety decision.

Caught 3 of 4 AI-labelled near-misses (blind, pre-registered labels); the engineer reviews 12 clips instead of 141 s of video / 51 candidates. Test split: 1 of 2 caught, 6 flags of 25 (split made after v1 saw all events; v3 designed from v1/v2 failures).

▶ Run the demoOpen the ledger🎮 Spot the near-miss

CloseCall finds possible near-misses in traffic-camera video for an engineer to review: a tracker measures every close pass, vision models triage the clips for an engineer to review, and the fix memo cites the labelled near-misses. Heads and plates are pixelated in tracked boxes; untracked people may not be. No identity is ever extracted.

AI-LABELLED NEAR-MISS · PET 0.13 s
A car → B person · Respubliki / Profsoyuznaya, Tyumen
Label: near-miss (AI, blind, pre-registered). Model (v3): flagged it · “The car travels straight through the intersection while the pedestrian crosses perpendicularly; neither vehicle nor pedestrian alters their speed or trajectory to avoid the other.”
vote: Qwen3.6-35B-A3B + MiniMax-M3 + gemma-4-31B-it · heads + plates pixelated in tracked boxes
Footage watched
141s
4 intersections · 592 tracks
Naive flags (PET < 4 s)
1165
what a pixel tracker alone reports
Plausible conflicts
51
crossing paths, both moving
Flagged for review
12
3 of the 4 AI-labelled near-misses are among them

Measured on this run (2026-10-02T19:50:33Z). Footage: CC BY-SA 4.0 public intersection clips from Wikimedia Commons, not a live city feed.

01 · The problem

Crash data only arrives after someone is hurt.

A city traffic engineer asking council to redesign an intersection “where nobody has died” has no evidence yet. The camera on the pole has recorded every close call for years, but nobody watches it: near-miss studies mean a person reviewing hours of video by hand, so they are rare and expensive.

02 · The solution

Measure every close call. Triage the likely ones.

A cheap first stage (YOLO26 tracking + post-encroachment time) flags every pair of road users that passed through the same spot within 4 seconds. Only those clips go to vision models, which reject most screen-space illusions and flag the likely conflicts, in plain words, for an engineer to review.

03 · Why it matters

Evidence before the funeral, not after.

Near-misses happen far more often than crashes, so they show a dangerous design in weeks instead of years. The ledger turns them into evidence a council can read: the clip, the timing, the explanation, the fix.

04 · How it’s used

Ask in plain words. Get the clips and the memo.

The engineer (or the consultancy doing the study) types “turning cars cutting across pedestrians”, gets the matching clips with timing and explanation, and exports a fix memo that cites every clip. Every step is traced in W&B Weave so the result can be audited.

How it works

YOLO26 + ByteTrack

Tracks every car, bus, truck, cyclist and pedestrian; computes post-encroachment time per pair.

NVIDIA Nemotron + embeddings

Nemotron-3 Omni was the video verifier in v1 and v2; nemotron-3-embed-1b powers plain-language search; the memo model is Nemotron-3.5.

VAST Data (adapter, stand-in)

Stand-in on this run: the VastStore adapter (clips, ledger rows, vectors, VSS ingest with a near-miss prompt) is written but not yet run against VAST.

W&B Inference + Weave

Runs the triage vote (Qwen3.6, MiniMax-M3, Gemma-4) on CoreWeave GPUs, drafts the memo, traces every call and scores each verifier against blind labels.

Measured precision of the verifier on this run: 33% versus 9% if every stage-1 flag were reported. See the eval, misses included →