Haga
Haga is the verification layer for physical AI: validated physics-consistency detectors for world-model outputs, and adversarial stress tests for the policies trained inside them. Reproducible numbers, shown failures — never a self-graded pass.
How it works
We apply one adversarial methodology to the two artifacts that matter — the generated world and the policy inside it. Every result ships with defined thresholds and shown failure cases.
Submit
Share a robot policy, world-model video, or both — plus your success criteria. We confirm scope and acceptance within 48 hours.
Verify artifacts
We parse the submitted artifacts and lock a reproducible run plan: paired seeds, fixed conditions, and measurable thresholds before execution.
Stress-test
Policy-stress probes score robustness under physics perturbations. World-model checks score physics-consistency with calibrated detectors.
Report outcomes
You get numeric reports with defined thresholds, shown failures, and enough detail for release reviews, safety cases, or investor diligence.
Problem
Manipulation benchmarks report 89.4% success in simulation; robots succeed in 12% of real household tasks (Stanford AI Index, 2026). That gap is where unverified physics lives. World-model builders and robot-learning labs self-report results; generated worlds ship without an independent physics check; policies that look robust inside a flawed simulation fail on deployment. Institutional leaderboards exist but are slow and inaccessible — there is no independent layer teams can buy rather than build.
Vision
Haga applies one adversarial methodology to both artifacts that matter: the generated world (is its physics consistent?) and the policy acting inside it (does it survive physical stress?). No competitor spans both. The output is always the same — reproducible numeric reports with defined thresholds and shown failure cases — built toward continuous scoring the industry can cite.
Trusted eval infrastructure that world-model builders, labs, and robotics companies rely on the way software teams rely on independent security audits — because self-reported physics stopped being credible.
Market
Capital and engineering effort are flooding robot-model builders, simulation tooling, and synthetic data — but independent physics-consistency checking across both world models and policies remains an open gap.
$27.6B
raised by physical AI and robotics startups across 1,009 deals in 2025 — more than double 2024.
$5.8B → $28.6B
AI world models market, 2025 to 2034 (58.2% CAGR). Interactive physics simulators — closest to Haga’s target environments — ~31.5% of that market.
Active peers
Antioch, Bifrost AI, Instance, and Robocurve are building simulation, synthetic data, and evaluation tooling in adjacent layers — the category is real and moving, but no one spans both world-model and policy verification.
TechCrunch / YC
Evidence
One adversarial methodology, two artifacts: the generated world and the policy inside it — with defined thresholds, paired seeds, and shown failure cases.
89.4% → 12%
sim benchmark success vs real household tasks (Stanford AI Index 2026)
1.000 recall
physics-violation detection, 0 false positives — calibrated checker v0
0% → 100%
Physics-IQ real quiet vs CogVideoX held-out (n=9, protocol v1) via static_hover
Roadmap
Both halves of the methodology exist and produce numbers today. Each next step extends a validated instrument — never a claim ahead of its data.
Stage 1
Tiered mass/friction stress on robosuite Lift, Stack, PickPlaceCan, and Door (50×4, Wilson CIs). Lift success degrades 1.00 → 0.26; Stack 0.96 → 0.20; PickPlaceCan 1.00 → 0.24; Door 0.60 → 0.52 (success-primary gate).
Stage 2
Calibrated detectors for teleportation, anti-gravity, causeless impulses, and interpenetration. Recall 1.000, zero false positives under tracking noise.
Stage 3
CoTracker3 + Physics-IQ Verified: real 0%, CogVideoX I2V cohort (n=6, seeds 0–1) 100% via static_hover freeze/hover failure. One documented generative failure mode — larger scenario sweeps still open. Cosmos/NIM deferred where geo-blocked.
Stage 4
Multi-task suites, methodology reports, then API scoring in customer release pipelines — the independent signal labs and world-model builders cite rather than build.
Team
Mushood Hanif is building Haga end-to-end — product, evaluation harness, and evidence. First hire / co-founder search targets ML systems and robotics eval depth; no multi-founder optics.
Mushood Hanif
Founder
Product, full-stack, and evaluation systems — Next.js, Python, sim-first physics verification; building Haga's harness, site, and evidence pack
Private evaluation
Tell us what you want scored — robot policy under physics stress, generative world-model video, or both. We acknowledge every submission within 48 hours.
48-hour acknowledgment SLA on every submission.
Read the public methodology — what we tested, how we score, and claim boundaries.