Haga — World-model validation for physical AI
logoHaga

Lab

How we employ Haga.

Running experiments with reproducible seeds, defined thresholds, and shown failure cases — the same adversarial methodology on world-model outputs and the policies trained inside them.

Each experiment below is a narrow, auditable instrument: one question, defined thresholds, and numeric reports with shown failure cases. We widen the suite only when the evidence already in hand supports it.

Methodologythresholds, seeds, claim boundaries, and links to technical specs.

Featured evidence

Tiered mass/friction stress on robosuite Lift: success degrades under severe stress, with grasp-slip as the dominant failure mode.

View full report

Charts

Success vs grasp-failure across severity

rate

Mean peak end-effector force

N (contact proxy)

World-model flag snapshot

static_hover 0% → 100%

Real quiet negative control 0%; Physics-IQ held-out cohort 100% via static_hover. One documented generative failure mode with paired seeds and claim boundaries.

Open Physics-IQ report