logoHaga

Blog

Practical guides on verification, benchmarking, and reproducible evaluation for physical AI — written for builders who need claim boundaries, not demo reels.

Need a private eval? Request one or subscribe via RSS.

Verification··3 min read

Physics-violation detection: current benchmark results

100% detection of injected physics violations at zero false flags on clean runs — with the full methodology, thresholds, and confidence intervals published.

Read article

Navigation··6 min read

Obstacle avoidance metrics beyond collision counts

Zero collisions can still mean a robot that freezes, brushes, or deadlocks. Here is the avoidance metric stack that distinguishes safe navigation from lucky navigation.

Read article

Physical AI··6 min read

Physical AI verification: the security-audit model

The software industry solved self-certification with independent audits. Physical AI verification should borrow that structure: separated checkers, frozen protocols, and adversarial evidence.

Read article

Manipulation··6 min read

Manipulation metrics that survive a skeptic

A skeptic asks: did the grasp hold, did placement land, did recovery terminate? Here is the manipulation metric stack that answers without being gamed.

Read article

Simulation··5 min read

MuJoCo vs robosuite vs Isaac Lab for evaluation

Training wants throughput; evaluation wants determinism and contact realism. A comparison of MuJoCo, robosuite, and Isaac Lab through the verification lens.

Read article

Evaluation··6 min read

Benchmark gaming in robot learning: when leaderboards lie

Every benchmark that matters gets gamed. Here are the failure modes of robotics leaderboards — and the protocol discipline that keeps numbers honest.

Read article

Simulation··5 min read

Domain randomization vs adversarial stress testing

Domain randomization is a training technique. Stress testing is an evaluation instrument. Using one to certify the other is a category error with real consequences.

Read article

World-model verification··5 min read

Generative video that looks right but breaks physics

Cinematic quality is not physical consistency. Here are the reproducible failure classes we see in generative video — and how to score them honestly.

Read article

World-model verification··5 min read

World model evaluation metrics: why FVD isn't physics

A world model can match video statistics and still freeze objects mid-air. Distributional metrics are weak instruments for physical understanding — here is what to measure instead.

Read article

Evaluation··5 min read

Sim-to-real benchmark design: measuring with scarce hardware

Benchmark design should make the evaluation protocol shareable and reproducible. Pass rates without failure modes are not diligence.

Read article

Physical AI··5 min read

Embodied AI trust gap: mixing demos, benchmarks, and claims

Physical-AI teams often blend sim success rates, world-model FID-style scores, and a manipulation demo into one headline. Independent evaluation separates them.

Read article

Navigation··5 min read

Mobile robot navigation evaluation: beyond success rate

Navigation stacks need evaluation that stresses localization drift, sensor dropout, and obstacle behavior — not only clean-lab success rates.

Read article

Simulation··4 min read

Robot simulator evaluation: verification over cinema

A simulator can be excellent for training and weak for release evaluation unless you apply an independent physics-consistency check with reproducible seeds.

Read article

Manipulation··5 min read

Pick-and-place verification: beyond the one-video demo

A successful demo does not prove a policy survives object mass changes, contact noise, or friction resampling. Here is how to verify pick-and-place claims.

Read article

Physical AI··5 min read

Why robot hands are hard: making progress measurable

A new dexterous-hand demo goes viral every few months. Without reproducible grasp protocols and failure taxonomies, progress stays unmeasurable.

Read article

World-model verification··5 min read

World-model evaluation: clips vs reproducible claims

If a world model looks realistic in a few clips, that is evidence — not proof. Evaluation turns visual plausibility into a reproducible, bounded claim.

Read article

Physical AI··5 min read

Sim-to-real trust gap: physics trust is the bottleneck

Higher simulator fidelity does not automatically create trust. Independent evaluation of physics assumptions — with shown failures — is what changes release decisions.

Read article

World-model verification··6 min read

Physics-IQ held-out cohort: real 0%, CogVideoX 100%

A fixed Physics-IQ verification protocol on a held-out generative cohort: zero flags on real quiet video, 100% flags on CogVideoX via static_hover — and what that does not prove.

Read article