Proof

Verifiable proof for every benchmark claim.

Start with the claim. Follow the evidence.

For executives, this is the trust path. For engineers, this is the audit path. Every headline number points back to scoped runs, artifacts, hashes, and the full registry.

Open the Evidence Registry
01 Claim

What is being asserted, where it applies, and where it does not.

02 Evidence

Canonical records carry artifacts, SHA256 hashes, Integrity Roots where available, runtime data, and labels.

03 Verification

Probe answers are checked against source records. Query cost is measured as archive depth grows.

01 What You Can Trust Here

The public claims on this site are tied to a canonical evidence set, not loose performance language.

  • 7 canonical runs are marked in the registry.
  • 24 total records document canonical, validation, regression, partial, and experimental work.
  • The largest completed continuous nuPlan run archived 2,000,000 telemetry records.
  • Canonical proof records show 0 integrity failures.
Inspect every run record →
02 What It Does Not Prove

It does not claim universal recall, production AV certification, model reasoning quality, or arbitrary dataset performance.

  • Partial runs are labeled as partial.
  • Regression and experimental runs are not elevated into the canonical claim.
  • Live demonstrations are separated from deterministic replay benchmarks.
Read the exact claim boundary →
03 Why This Page Exists

The registry is dense. This page gives decision makers the short proof path, then sends technical reviewers to the raw records.

  • For a technical audit, open the registry.
  • For scope and language, open the claim page.
  • For reproduction flow, open the benchmark page.
See methodology and reproduction →