logoHaga

Lab

How we employ Haga.

Running experiments with reproducible seeds, defined thresholds, and shown failure cases — the same adversarial methodology on world-model outputs and the policies trained inside them.

Each experiment below is a narrow, auditable instrument: one question, defined thresholds, and numeric reports with shown failure cases. We widen the suite only when the evidence already in hand supports it.

Methodology — thresholds, seeds, claim boundaries, and links to technical specs.

Featured evidence

Tiered mass/friction stress on robosuite Lift: success degrades under severe stress, with grasp-slip as the dominant failure mode.

View full report

Charts

Success vs grasp-failure across severity

rate

Mean peak end-effector force

N (contact proxy)

World-model flag snapshot

static_hover 0% → 100%

Real quiet negative control 0%; Physics-IQ held-out cohort 100% via static_hover. One documented generative failure mode with paired seeds and claim boundaries.

Open Physics-IQ report