logoHaga

FAQ

Questions about evaluation, scope, and pilot engagement.

These are the questions we hear most often from labs, physical-AI teams, and investors doing diligence on model or simulator claims.

How is Haga different from a simulator or synthetic-data platform?

Simulators generate environments; Haga audits them. Haga does not build physics simulators or synthetic datasets. It checks whether generated worlds and policies inside them are physics-consistent under reproducible adversarial stress — and returns numeric reports with shown failures.

What do you mean by independent verification for physical AI?

Independent means the checker is not the same team that produced the artifact. For world models and robot policies, that separation matters: self-reported benchmarks can improve when the test itself changes. Haga's role is to be the outside check that releases, investors, or labs can cite.

What artifacts can you evaluate?

We accept robot policies, world-model or generative-video outputs, or both. If you are unsure which is the right starting point, choose exploratory intake and we will propose the narrower, evidence-supported path.

What is the evaluation SLA?

We acknowledge every submission within 48 hours and return a structured report relative to scope. Larger suites and multi-task sweeps are scheduled as pilots with explicit timeline gates.

Can you evaluate hardware or real-robot deployments directly?

Current verification is simulation-first. That is intentional: sim evidence is reproducible, repeatable, and auditable at lower cost before hardware validation. We explicitly label claim boundaries so hardware-in-the-loop differences are not hidden.

Do you only check world models, or do you also audit robot policies?

Both. Pillar 1 is policy-stress evaluation under physics perturbations. Pillar 2 is world-model physics-consistency checking. The value is that both artifacts use the same adversarial methodology and the same complementary evidence standard.

How do you handle private or pre-release model submissions?

Eval intake is private by default. We do not publish client artifacts, unreleased weights, or internal datasets. Public reporting is limited to aggregated methodology and anonymized cohort evidence, unless you opt in to named attribution.

Who should submit an eval — labs, startups, or investors?

Labs and model teams can use evals for release readiness or internal QA. Investors can use them as diligence artifacts when evaluating portfolio companies. If the goal is external validation, the right intake is a scoped pilot with explicit report ownership.

Why is the founder credibility relevant to evaluation credibility?

Evaluation credibility depends on reproducibility and claim discipline, not team size. Haga ships full methodology, threshold definitions, seeds, and failure cases. Solo-founder status reduces coordination overhead and keeps the evidence chain short.

Can I see example reports before submitting?

Yes. Public method summaries and live lab experiments are the best example of artifact format. For private examples, request a short methodology brief during intake and we will share redacted report structure.