Methodology
Inspect what we tested and how we scored it.
Haga’s public posture: methodology, detector definitions, and Lab evidence on this site; customer artifacts and comparative engagement data stay private. Same adversarial approach on two artifacts — the policy under physics stress, and the generated world it trains in.
Wedge
Fast, private, sim-first physics verification for robot policies and generative world-model outputs — not a public leaderboard, not a sim platform, not a training loop.
Two pillars
Pillar 1 — Policy stress
Adversarial mass/friction tiers on robosuite Lift, Stack, PickPlaceCan, and Door. Success rates with Wilson CIs, grasp/place failure where defined, and shown severe-tier failures — never a gate-only PASS.
Pillar 2 — World-model physics
Calibrated physics-violation detectors (teleport, anti-gravity, impulse, interpenetration), then CoTracker3 → VIDEO_CHECKS on real Physics-IQ and CogVideoX I2V. Real cohort quiet (0%); discovery n=6 post-hoc via static_hover; held-out protocol v1 n=9 flag rate 1.000 (Wilson [0.701, 1.000]) — confirmatory, not pooled with discovery. Not Cosmos / NIM.
Claim boundaries
- Sim-first policy stress — not a real-robot leaderboard.
- CogVideoX I2V cohort scale (n=6, seeds 0–1); larger scenario sweeps before broad generative claims.
- Never relabel CogVideoX results as Cosmos / Genie / NIM.
- Gate-tier PASS is incomplete without severe-tier degradation and shown failures.
Full write-up with citations and claim boundaries: Article — State of Sim Physics Consistency, v1. Numeric charts: Lab. Detailed methodology narrative: this page.