Evaluation · W&B Weave

Does the triage catch the near-misses, and how much review does it save?

The scoring rule was committed to the repo (eval/PREREGISTRATION.md) before any label was written, and the labels (eval/labels.json) were committed before any verifier ran on the set. Each verifier version has its own pre-registration commit and was run once; none of the numbers below was re-run to improve it.

Caught 3 of 4 AI-labelled near-misses (blind, pre-registered labels); the engineer reviews 12 clips instead of 141 s of video / 51 candidates. Test split: 1 of 2 caught, 6 flags of 25 (split made after v1 saw all events; v3 designed from v1/v2 failures).

Every pre-registered run · same 51 events · same blind labels (unsure excluded)
VerifierSetCaught (recall)Flags (review load)Labelled not a near-miss (FP)PrecisionPre-registration
v1 · Nemotron Omni (39) + Llama-3.2 fallback (12), whole clipall 510 / 49 of 44 scored90%PREREGISTRATION.md
v2 · Nemotron Omni, zoomed clip + rule (2 calls errored → not flagged)all 512 / 410 of 44 scored820%PREREGISTRATION-v2.md
v3 · storyboard, 2-of-3 votedev (chosen here: optimistic)2 / 26 of 26 candidates340%PREREGISTRATION-v3-split.md
v3 · storyboard, 2-of-3 voteTEST (held out, one run)1 / 26 of 25 candidates325%PREREGISTRATION-v3-test.md
v3 · storyboard, 2-of-3 voteall 51 (includes dev)3 / 412 of 51 candidates633%PREREGISTRATION-v3-test.md

Each version was written after the previous one failed and committed before it ran. v3 was chosen on the dev split (recall first, then fewest flags) and run once on test. Test has only 2 labelled near-misses, so its recall is a coarse measure. The split was made after v1 saw all 51 events, and v3’s design was informed by v1/v2 failures, so test is not fully held out. The labels were never changed. The ledger uses v3 · storyboard · 2-of-3 vote (Qwen3.6-35B, MiniMax-M3, Gemma-4-31B on W&B Inference).

Ledger verifier (v3), all 51: precision
33%
3 of 9 flags correct
Recall
75%
3 of 4 labelled near-misses found
Stage 1 alone (baseline)
9%
precision if every PET flag were reported
Scored / excluded
44 / 7
events scored / labelled “unsure”
True positives · 3
  • tyumen-prof-397-437 car then person: model said “The car travels straight through the intersection while the pedestrian crosses perpendicularly; neither vehicle nor pedestrian alters their speed or trajectory to avoid the other.”
    label note: car turns through crosswalk; pedestrian steps onto the spot 0.13 s later
  • tyumen-prof-9-188 person then car: model said “Pedestrian A crosses the road on a crosswalk while car B drives through the intersection; A clears the crossing point just 0.2s before B arrives, leaving them within a meter of each other, but neither brakes nor swerves.”
    label note: car passes the pedestrian's curb-edge position 0.2 s after
  • tyumen-prof-437-567 person then car: model said “The pedestrian crosses the street from left to right while the car drives straight through the intersection from right to left, passing very close to the pedestrian's path without either yielding or braking.”
    label note: turning car crosses the crosswalk 0.6 s after the pedestrian
False positives (model said near-miss, label said no) · 6
  • pieix-1066-410 bicycle then bus: model said “The cyclist and bus are both on the road at the same depth, and their paths cross at the red circle. The cyclist passes through the intersection without stopping or swerving, and the bus continues moving forward.”
  • tyumen-prof-310-157 car then car: model said “The blue car (A) turns left across the path of the silver car (B), which is driving straight. They pass each other at the intersection without either vehicle braking or swerving.”
  • tyumen-prof-486-407 car then car: model said “Car A turns left across the path of Car B, which is proceeding straight. They pass each other at the intersection without either vehicle braking or swerving abruptly.”
  • tyumen-prof-30-188 car then car: model said “Car A turns left across the path of Car B, which is traveling straight; their ground paths intersect and they pass within a car length of each other without either braking or swerving.”
  • chiangmai-219-352 car then car: model said “The silver pickup truck (A) travels straight through the intersection while the dark sedan (B) turns right across its path. They pass each other at the same depth without either vehicle braking or swerving.”
  • chiangmai-115-176 motorcycle then motorcycle: model said “Both motorcycles are traveling on the road and their paths cross at the intersection; they pass each other closely without any sudden braking or swerving.”
Missed (label said near-miss, model said no) · 1
  • tyumen-vodo-3-38 car then person: model said “The car (A) travels along the main road while the cyclist (B) crosses on a perpendicular path in the background; they are at significantly different distances from the camera and do not come close to each other.”
    label note: pedestrian enters the crossing 0.87 s after the car passes

Flagged by the model but labelled “unsure” (not scored): pieix-958-1066, tyumen-prof-407-732, tyumen-prof-654-786

Limits, stated plainly

  • Small sample: 51 events from 141 s of public footage. Indicative, not a benchmark.
  • Labeller: Claude Code (builder), blind to verifier verdicts, from 3 anonymised keyframes per event. Not a traffic engineer, and the same agent that built the pipeline.
  • PET is measured in screen pixels without a ground-plane calibration, which is why stage 1 alone is so noisy.
  • v3 (the ledger): 0 of 51 verdicts from a fallback model. v1: 12 of 51 came from the Llama 3.2 Vision fallback on one keyframe; v2: 2 of 51 calls errored and count as not flagged.

Open traces + evaluation in Weave ↗