logoHaga

Methodology

Inspect what we tested and how we scored it.

Haga’s public posture: methodology, detector definitions, and Lab evidence on this site; customer artifacts and comparative engagement data stay private. Same adversarial approach on two artifacts — the policy under physics stress, and the generated world it trains in.

Wedge

Fast, private, sim-first physics verification for robot policies and generative world-model outputs — not a public leaderboard, not a sim platform, not a training loop.

Two pillars

  • Pillar 1 — Policy stress

    Adversarial mass/friction tiers on robosuite Lift, Stack, PickPlaceCan, and Door. Success rates with Wilson CIs, grasp/place failure where defined, and shown severe-tier failures — never a gate-only PASS.

    Lab results

  • Pillar 2 — World-model physics

    Calibrated physics-violation detectors (teleport, anti-gravity, impulse, interpenetration), then CoTracker3 → VIDEO_CHECKS on real Physics-IQ and CogVideoX I2V. Real cohort quiet (0%); discovery n=6 post-hoc via static_hover; held-out protocol v1 n=9 flag rate 1.000 (Wilson [0.701, 1.000]) — confirmatory, not pooled with discovery. Not Cosmos / NIM.

    Lab results

Claim boundaries

  • Sim-first policy stress — not a real-robot leaderboard.
  • CogVideoX I2V cohort scale (n=6, seeds 0–1); larger scenario sweeps before broad generative claims.
  • Never relabel CogVideoX results as Cosmos / Genie / NIM.
  • Gate-tier PASS is incomplete without severe-tier degradation and shown failures.

Full write-up with citations and claim boundaries: Article — State of Sim Physics Consistency, v1. Numeric charts: Lab. Detailed methodology narrative: this page.