Lab
Running experiments with reproducible seeds, defined thresholds, and shown failure cases — the same adversarial methodology on world-model outputs and the policies trained inside them.
Each experiment below is a narrow, auditable instrument: one question, defined thresholds, and numeric reports with shown failure cases. We widen the suite only when the evidence already in hand supports it.
Methodology — thresholds, seeds, claim boundaries, and links to technical specs.
Featured evidence
Tiered mass/friction stress on robosuite Lift: success degrades under severe stress, with grasp-slip as the dominant failure mode.
Success vs grasp-failure across severity
rate
Mean peak end-effector force
N (contact proxy)
World-model flag snapshot
static_hover 0% → 100%
Real quiet negative control 0%; Physics-IQ held-out cohort 100% via static_hover. One documented generative failure mode with paired seeds and claim boundaries.
Policy verification
ShippedTiered mass/friction stress on robosuite Lift: success degrades under severe stress, with grasp-slip as the dominant failure mode.
1.00 → 0.26
success rate, nominal → severe
74%
severe episodes: grasp acquired, then lost
Policy verification
ShippedTiered mass/friction stress on robosuite Stack: success degrades under severe stress, with grasp-slip as the dominant failure mode.
0.96 → 0.20
success rate, nominal → severe
80%
severe episodes: grasp acquired, then lost
Policy verification
ShippedTiered mass/friction stress on robosuite PickPlaceCan: success degrades under severe stress, with grasp-slip as the dominant failure mode.
1.00 → 0.24
success rate, nominal → severe
76%
severe episodes: grasp acquired, then lost
Policy verification
ShippedPanel-scaled mass/friction stress on robosuite Door with a success-primary gate: shallower degradation curve, door-did-not-open failures documented.
0.60 → 0.52
success rate, nominal → severe
48%
severe episodes: door did not open
World-model checker v0
ShippedCalibrated physics-violation detector: 1.000 recall on 399 violated trajectories, 0 false positives on 200 negatives.
1.000
recall on 399 violated trajectories
0
false positives on 200 negatives
Tracking + CogVideoX cohort
ShippedFrozen VIDEO_CHECKS: real quiet (0%); discovery n=6 post-hoc 100%; held-out protocol v1 n=9 100% [0.701, 1.000] — all via `static_hover`.
0%
real negative control
100%
held-out flag rate (n=9 [0.701, 1.000])