Technical article · 2026-07-19
State of Sim Physics Consistency in Open Manipulation Policies, v1
What we tested, how we scored it, and what failed. Methodology, detector definitions, and Lab evidence are the public trust signal. Partner artifacts and comparative data stay private. Published numbers live in the Lab.
Lab chartsBounded demoMethodology summary
Wedge
Fast, private, sim-first physics verification for robot policies and generative world-model outputs — not a public leaderboard, not a sim platform, not a training loop.
Why this matters
Stanford HAI’s 2026 AI Index documents a familiar gap: robotic manipulation reaches 89.4% success in software sims (RLBench) but only 12% on real household tasks. That gap is where unverified physics lives.
Institutional leaderboards occupy a different tier — large-scale, real-robot, slow-cycle. Haga targets the pre-deployment, commercially accessible layer: private stress reports with defined thresholds and shown failure cases. Physical AI / robotics startups raised $27.6B across 1,009 deals in 2025 — capital into systems that need verification faster than into verification itself.
Claim under test
Does a deterministic OSC_POSE baseline stay physically consistent when object / panel mass and contact friction are randomized within documented severity tiers? And can the same adversarial posture detect physics violations in generative video?
- Claim: sim-first degradation curves with Wilson CIs, documented failure modes, and a calibrated video checker that fires on CogVideoX freeze/hover.
- Not claimed: real-hardware validation; Cosmos / Genie / NIM scores; that mild-gate PASS equals deployment readiness; broad generative quality from a 3-clip pilot.
Pillar 1 — Policy stress
Robosuite on MuJoCo · Panda · OSC_POSE · 50 episodes per condition · paired seeds 0..49 · Wilson 95% CIs · mild / moderate / severe mass+friction tiers. Door uses a success-primary gate (grasp N/A for rotating handle).
| Task | Nominal → severe | Mild gate | Severe failure |
|---|---|---|---|
| Lift | 1.00 → 0.26 | PASS | Grasp slip 0.74 |
| Stack | 0.96 → 0.20 | PASS (≈0.85) | Grasp fail 0.80 |
| PickPlaceCan | 1.00 → 0.24 | PASS | Grasp slip 0.76 |
| Door | 0.60 → 0.52 | PASS (success-primary) | Door did not open 0.48 |
Mild-gate PASS is the entry bar. The diligence story is the degradation curve and shown failures — not a single green badge. Published headline rates are in the Lab (sanitized metrics JSON).
Pillar 2 — Checker + generative video
Position-only detectors (permanence, ballistic, contact) calibrated on MuJoCo ground truth: negatives flag at 0.00; teleport / hover / impulse / penetration at 1.00 (recall 1.000, FPR 0.000).
Video path: RGB → CoTracker3 → VIDEO_CHECKS including static_hover. Real Physics-IQ negative control: flag rate 0.000. Discovery CogVideoX cohort (n=6, seeds 0–1): flag rate 1.000 via static_hover — post-hoc. Not Cosmos / NIM.
Held-out generative validation (protocol v1)
Question. After freezing VIDEO_CHECKS / static_hover, does freeze/hover still fire on a pre-registered CogVideoX matrix never used to tune the detector?
Method. Thresholds locked 2026-07-19. Model THUDM/CogVideoX-5b-I2V ; scenarios 0001–0003 × seeds {2, 3, 4} (n=9), disjoint from discovery seeds 0–1.
Results. Held-out flag rate 1.000 (9/9) · Wilson 95% CI [0.701, 1.000] — all static_hover. Real control remains quiet. Discovery stays labeled separately.
Limits. One scene family, one failure mode, one open I2V model — not a broad generative-quality verdict and not Cosmos / NIM.
Request. Critique the detector, or send a public artifact with provenance for a free bounded evaluation.
Charts: Lab · Physics-IQ · Bounded demo.
Competitive whitespace
RoboArena is institutional real-robot eval. Robocurve emphasizes real-hardware benchmarks. Instance targets physics-aware QA for AI-generated video. Haga’s wedge is both policy stress and generative physics under one private-evaluation product — sim-first, fast turnaround, shown failures — without pretending to replace institutional leaderboards.
What to do next
- Inspect the Lab, the bounded demo, and the claim boundaries above.
- Critique the held-out protocol or send a public artifact for a free bounded evaluation.
- Submit for private evaluation — artifact type, timeline, company; we acknowledge within 48 hours.
- Expect a scoped report: numeric tables, CIs, failure cases, limitations — not a marketing PASS badge.
Citations
- Stanford HAI, 2026 AI Index Report — Technical Performance
- RoboArena coverage (The Next Web, June 2026)
- Physical AI Startup Statistics (Mean CEO, May 2026)
- Robocurve (Y Combinator)
Figures as of 2026-07-19. Prefer published Lab metrics if schema dates diverge. Sim-only limitation applies to all Pillar 1 numbers. Public methodology and Lab evidence are the trust signal; private evaluation engagements remain the product.