# Haga — World-model validation for physical AI > Independent stress-testing of simulated worlds and robot policies before they ship into the real one. Reproducible seeds, defined thresholds, shown failure cases. ## Context World-model validation for physical AI — independent stress-testing of simulated worlds and robot policies. Haga is an independent verification layer for physical AI. Two complementary instruments: - **Pillar 1 — policy verification:** tiered mass/friction physics perturbations on robot manipulation policies (Lift, Stack, PickPlaceCan, Door) with gate thresholds, Wilson confidence intervals, and logged failure cases. - **Pillar 2 — world-model physics:** calibrated physics-violation detectors (teleportation, anti-gravity, impulses, freezes) applied to generative-video and real-video cohorts, including held-out Physics-IQ evaluations. All methodology, thresholds, seeds, and failure cases are public. Partner artifacts and comparative data stay private. ## Lab — live experiments - [Lift stress](https://haga.mushoodhanif.com/lab/lift-stress/) — tiered mass/friction perturbations on the Lift task - [Stack stress](https://haga.mushoodhanif.com/lab/stack-stress/) — tiered mass/friction perturbations on the Stack task - [PickPlaceCan stress](https://haga.mushoodhanif.com/lab/pickplacecan-stress/) — tiered mass/friction perturbations on PickPlaceCan - [Door stress](https://haga.mushoodhanif.com/lab/door-stress/) — tiered mass/friction perturbations on Door - [Physics checker](https://haga.mushoodhanif.com/lab/physics-checker/) — recall/false-positive calibration of physics-violation detectors - [Physics-IQ](https://haga.mushoodhanif.com/lab/physicsiq/) — held-out real-video controls vs generative-video cohorts ## Blog - [Physics-violation detection: current benchmark results](https://haga.mushoodhanif.com/blog/physics-violation-detection-benchmark-results) — Machine-generated results from the Haga physics checker and policy stress gates. Every number on this page is traceable to a JSON artifact in the haga-core repository. (markdown: [Physics-violation detection: current benchmark results](https://haga.mushoodhanif.com/blog/physics-violation-detection-benchmark-results.md)) - [Obstacle avoidance metrics beyond collision counts](https://haga.mushoodhanif.com/blog/obstacle-avoidance-evaluation-metrics) — Collision counts miss near-misses, deadlock, and failure mode. Mobile-robot avoidance evaluation needs proximity, availability, and recovery metrics. (markdown: [Obstacle avoidance metrics beyond collision counts](https://haga.mushoodhanif.com/blog/obstacle-avoidance-evaluation-metrics.md)) - [Physical AI verification: the security-audit model](https://haga.mushoodhanif.com/blog/physical-ai-independent-verification) — Self-reported benchmarks are not diligence. Independent checkers, frozen protocols, and adversarial testing: the audit template for physical AI. (markdown: [Physical AI verification: the security-audit model](https://haga.mushoodhanif.com/blog/physical-ai-independent-verification.md)) - [Manipulation metrics that survive a skeptic](https://haga.mushoodhanif.com/blog/manipulation-evaluation-metrics) — Grasp success rate alone hides slip, regrasp, and placement errors. Skeptic-proof manipulation evaluation splits motion, grasp, and placement evidence. (markdown: [Manipulation metrics that survive a skeptic](https://haga.mushoodhanif.com/blog/manipulation-evaluation-metrics.md)) - [MuJoCo vs robosuite vs Isaac Lab for evaluation](https://haga.mushoodhanif.com/blog/robot-simulator-comparison) — Simulators differ in physics fidelity and repeatability. For verification, contact resolution and seed discipline matter more than visuals. (markdown: [MuJoCo vs robosuite vs Isaac Lab for evaluation](https://haga.mushoodhanif.com/blog/robot-simulator-comparison.md)) - [Benchmark gaming in robot learning: when leaderboards lie](https://haga.mushoodhanif.com/blog/robot-benchmark-gaming) — Benchmarks get gamed wherever they gate funding or attention. Overfitting, selective reporting, and protocol drift make leaderboards unreliable signals. (markdown: [Benchmark gaming in robot learning: when leaderboards lie](https://haga.mushoodhanif.com/blog/robot-benchmark-gaming.md)) - [Domain randomization vs adversarial stress testing](https://haga.mushoodhanif.com/blog/domain-randomization-vs-stress-testing) — Randomized training domains improve transfer; adversarial stress testing measures robustness. Mixing them up is how robustness claims get inflated. (markdown: [Domain randomization vs adversarial stress testing](https://haga.mushoodhanif.com/blog/domain-randomization-vs-stress-testing.md)) - [Generative video that looks right but breaks physics](https://haga.mushoodhanif.com/blog/generative-video-physics-failures) — Sora, Veo, and CogVideoX outputs look cinematic and violate physics in reproducible ways. Failure-mode catalogs turn demos into comparable evidence. (markdown: [Generative video that looks right but breaks physics](https://haga.mushoodhanif.com/blog/generative-video-physics-failures.md)) - [World model evaluation metrics: why FVD isn't physics](https://haga.mushoodhanif.com/blog/world-model-evaluation-metrics) — FVD measures distributional similarity, not physical correctness. Trajectory-level checks and violation detectors are the metrics that survive scrutiny. (markdown: [World model evaluation metrics: why FVD isn't physics](https://haga.mushoodhanif.com/blog/world-model-evaluation-metrics.md)) - [Sim-to-real benchmark design: measuring with scarce hardware](https://haga.mushoodhanif.com/blog/sim-to-real-benchmark) — A credible sim-to-real benchmark needs protocol discipline, claim boundaries, and an explicit failure taxonomy — not just a leaderboard of successes. (markdown: [Sim-to-real benchmark design: measuring with scarce hardware](https://haga.mushoodhanif.com/blog/sim-to-real-benchmark.md)) - [Embodied AI trust gap: mixing demos, benchmarks, and claims](https://haga.mushoodhanif.com/blog/embodied-ai-trust-gap) — Embodied-AI marketing outpaces measurement. Separate policy robustness, world-model consistency, and hardware evidence before you call a system deployable. (markdown: [Embodied AI trust gap: mixing demos, benchmarks, and claims](https://haga.mushoodhanif.com/blog/embodied-ai-trust-gap.md)) - [Mobile robot navigation evaluation: beyond success rate](https://haga.mushoodhanif.com/blog/mobile-robot-navigation-evaluation) — Navigation can succeed in a clean map and hide weak localization and brittle obstacle handling. Probe failure modes before you claim robustness. (markdown: [Mobile robot navigation evaluation: beyond success rate](https://haga.mushoodhanif.com/blog/mobile-robot-navigation-evaluation.md)) - [Robot simulator evaluation: verification over cinema](https://haga.mushoodhanif.com/blog/robot-simulator-evaluation) — Simulators differ in physics fidelity, sensor modeling, and repeatability. The right choice for evaluation is not always the most cinematic one. (markdown: [Robot simulator evaluation: verification over cinema](https://haga.mushoodhanif.com/blog/robot-simulator-evaluation.md)) - [Pick-and-place verification: beyond the one-video demo](https://haga.mushoodhanif.com/blog/pick-and-place-verification) — Automated pick and place needs reproducible policy evaluation under physics stress — not a single successful rollout video. (markdown: [Pick-and-place verification: beyond the one-video demo](https://haga.mushoodhanif.com/blog/pick-and-place-verification.md)) - [Why robot hands are hard: making progress measurable](https://haga.mushoodhanif.com/blog/why-robot-hands-are-hard) — Dexterous robot hands lack a shared measurement standard beyond the demo reel. Independent evaluation turns highlight clips into comparable evidence. (markdown: [Why robot hands are hard: making progress measurable](https://haga.mushoodhanif.com/blog/why-robot-hands-are-hard.md)) - [World-model evaluation: clips vs reproducible claims](https://haga.mushoodhanif.com/blog/world-model-evaluation) — Realistic video is not proof of physical correctness. Independent world-model evaluation measures consistency, failure modes, and claim boundaries. (markdown: [World-model evaluation: clips vs reproducible claims](https://haga.mushoodhanif.com/blog/world-model-evaluation.md)) - [Sim-to-real trust gap: physics trust is the bottleneck](https://haga.mushoodhanif.com/blog/sim-to-real-trust-gap) — The sim-to-real gap is a trust problem, not only a hardware problem. Independent physics evaluation changes what you can claim before deployment. (markdown: [Sim-to-real trust gap: physics trust is the bottleneck](https://haga.mushoodhanif.com/blog/sim-to-real-trust-gap.md)) - [Physics-IQ held-out cohort: real 0%, CogVideoX 100%](https://haga.mushoodhanif.com/blog/physics-iq-held-out-cohort) — Protocol v1 held-out results on Physics-IQ: real quiet controls stay quiet, CogVideoX I2V fails uniformly via static_hover — with claim boundaries. (markdown: [Physics-IQ held-out cohort: real 0%, CogVideoX 100%](https://haga.mushoodhanif.com/blog/physics-iq-held-out-cohort.md)) ## Key pages - [Home](https://haga.mushoodhanif.com/) — positioning and contact - [Methodology](https://haga.mushoodhanif.com/methodology/) — public methodology and claim boundaries - [Lab overview](https://haga.mushoodhanif.com/lab/) — experiment hub - [Blog index](https://haga.mushoodhanif.com/blog/) — all posts - [FAQ](https://haga.mushoodhanif.com/faq/) — evaluation scope, SLA, and pilot engagement - [Demo](https://haga.mushoodhanif.com/demo/) — bounded, canned-scored demo - [Technical article](https://haga.mushoodhanif.com/article/sim-physics-consistency-v1/) — sim physics consistency of open manipulation policies, v1 - [Deck](https://haga.mushoodhanif.com/deck/) — investor deck - [About](https://haga.mushoodhanif.com/about/) — company and team