logoHaga
logo

Haga

World models are training tomorrow's robots. Nobody independently checks their physics.

Haga is the verification layer for physical AI: validated physics-consistency detectors for world-model outputs, and adversarial stress tests for the policies trained inside them. Reproducible numbers, shown failures — never a self-graded pass.

How it works

Four steps from artifact to reproducible report.

We apply one adversarial methodology to the two artifacts that matter — the generated world and the policy inside it. Every result ships with defined thresholds and shown failure cases.

  1. 1

    Submit

    Share a robot policy, world-model video, or both — plus your success criteria. We confirm scope and acceptance within 48 hours.

  2. 2

    Verify artifacts

    We parse the submitted artifacts and lock a reproducible run plan: paired seeds, fixed conditions, and measurable thresholds before execution.

  3. 3

    Stress-test

    Policy-stress probes score robustness under physics perturbations. World-model checks score physics-consistency with calibrated detectors.

  4. 4

    Report outcomes

    You get numeric reports with defined thresholds, shown failures, and enough detail for release reviews, safety cases, or investor diligence.

Problem

A flawed world teaches an agent the wrong physics. Both ship unchecked.

Manipulation benchmarks report 89.4% success in simulation; robots succeed in 12% of real household tasks (Stanford AI Index, 2026). That gap is where unverified physics lives. World-model builders and robot-learning labs self-report results; generated worlds ship without an independent physics check; policies that look robust inside a flawed simulation fail on deployment. Institutional leaderboards exist but are slow and inaccessible — there is no independent layer teams can buy rather than build.

Vision

The trust layer for the world-model era.

Haga applies one adversarial methodology to both artifacts that matter: the generated world (is its physics consistent?) and the policy acting inside it (does it survive physical stress?). No competitor spans both. The output is always the same — reproducible numeric reports with defined thresholds and shown failure cases — built toward continuous scoring the industry can cite.

Trusted eval infrastructure that world-model builders, labs, and robotics companies rely on the way software teams rely on independent security audits — because self-reported physics stopped being credible.

Market

Physical AI is scaling fast. Verification is underserved.

Capital and engineering effort are flooding robot-model builders, simulation tooling, and synthetic data — but independent physics-consistency checking across both world models and policies remains an open gap.

Evidence

Both pillars have running code and real numbers.

One adversarial methodology, two artifacts: the generated world and the policy inside it — with defined thresholds, paired seeds, and shown failure cases.

sim benchmark success vs real household tasks (Stanford AI Index 2026)

89.4% → 12%

sim benchmark success vs real household tasks (Stanford AI Index 2026)

physics-violation detection, 0 false positives — calibrated checker v0

1.000 recall

physics-violation detection, 0 false positives — calibrated checker v0

Physics-IQ real quiet vs CogVideoX held-out (n=9, protocol v1) via static_hover

0% → 100%

Physics-IQ real quiet vs CogVideoX held-out (n=9, protocol v1) via static_hover

Roadmap

Two pillars shipped narrow. Widening on evidence, not ambition.

Both halves of the methodology exist and produce numbers today. Each next step extends a validated instrument — never a claim ahead of its data.

  1. Stage 1

    Shipped — policy stress benchmark

    Tiered mass/friction stress on robosuite Lift, Stack, PickPlaceCan, and Door (50×4, Wilson CIs). Lift success degrades 1.00 → 0.26; Stack 0.96 → 0.20; PickPlaceCan 1.00 → 0.24; Door 0.60 → 0.52 (success-primary gate).

  2. Stage 2

    Shipped — physics-consistency checker v0

    Calibrated detectors for teleportation, anti-gravity, causeless impulses, and interpenetration. Recall 1.000, zero false positives under tracking noise.

  3. Stage 3

    Shipped — real-video + CogVideoX multi-seed cohort

    CoTracker3 + Physics-IQ Verified: real 0%, CogVideoX I2V cohort (n=6, seeds 0–1) 100% via static_hover freeze/hover failure. One documented generative failure mode — larger scenario sweeps still open. Cosmos/NIM deferred where geo-blocked.

  4. Stage 4

    Platform — continuous scoring as infrastructure

    Multi-task suites, methodology reports, then API scoring in customer release pipelines — the independent signal labs and world-model builders cite rather than build.

Monetization path (intent)

  • Benchmarking-as-a-service for design partners
  • Enterprise eval contracts for labs and physical-AI teams
  • API scoring for continuous world-model and policy checks

Team

Sole founder shipping the verification layer.

Mushood Hanif is building Haga end-to-end — product, evaluation harness, and evidence. First hire / co-founder search targets ML systems and robotics eval depth; no multi-founder optics.

Private evaluation

Submit your policy or world model for private evaluation.

Tell us what you want scored — robot policy under physics stress, generative world-model video, or both. We acknowledge every submission within 48 hours.

48-hour acknowledgment SLA on every submission.

Read the public methodologywhat we tested, how we score, and claim boundaries.