Real nuPlan Mini AV telemetry. 10 completed runs across 5 scale tiers. 3 canonical reference runs with full artifact verification. Primary benchmark domain.
Dataset
nuPlan is a large-scale AV planning dataset from Motional / nuTonomy. The Mini subset used in TeleMemetry benchmarks contains real nuPlan telemetry replayed from stored files as a streaming feed. It is not live vehicle data, not an NVIDIA DRIVE integration, and not a production safety validation.
Data lane: REAL NUPLAN SUMMARY. No raw telemetry frames are included in public artifacts. Numbers in result cards come from the run summary JSON and were not adjusted or normalized after the fact.
Dataset not included in public artifacts. Independent reviewers may request access to the dataset configuration and replay harness at AIFSOfficial@proton.me.
Canonical reference runs
| Run | Date | Records | Probes | Recall | Tok/Q | Runtime | RPS | GPU W avg | VRAM peak | SHA256 (prefix) | Integrity Root (prefix) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| REF Ref 001 | 2026-06-29 | 1,000,000 | 100 / 100 | 100% | 397 | 6h 39m 31s | 41.72 | 16.51W | 4.73 GB | 5a06e018… | 8ee1a1bf… |
| REF Ref 002 | 2026-06-30 | 1,000,000 | 500 / 500 | 100% | 397 | 11h 29m 55s | 24.16 | 25.32W | 4.895 GB | 37735071… | bd2e0eec… |
| REF Ref 003 | 2026-07-03 | 1,000,000 | 500 / 500 | 100% | 397 | 11h 55m 8s | 23.31 | 24.91W | n/a | 2e8c0f3e… | n/a |
Initial 1M run. 94/100 answer-layer citation accuracy (retrieval 100/100). The 6 misses were at the answer/citation layer, not the retrieval layer. SHA256, Integrity Root, and GPU CSV verified locally.
1M records, 500/500 retrieval recall. Second artifact also produced: gatekeeper-av-nuplan-1m-500q-rtx4000ada20gb-qwen25-7b-proofstress-20260630-131945.tgz (SHA256: c24bd4834f87843e730691cb3ff1841bef1bb0e605cdbf56fc2718bf7273999d). GPU CSV and Integrity Root verified locally.
Endurance run
| Run | Date | Records | Probes | Recall | Runtime | RPS | Replay tokens avoided (est.) | SHA256 (prefix) | Integrity Root (prefix) |
|---|---|---|---|---|---|---|---|---|---|
| END Endurance 001 | 2026-07-01 | 2,000,000 | 500 / 500 | 100% | 33h 38m | 16.51 | ~92,000,000 | 4e6db889… | ec0cbf43… |
2M-record continuous run. Largest completed nuPlan run. At a hypothetical full-replay baseline of ~100 tokens/record, the 2M archive would require ~200M context tokens per probe. TeleMemetry used ~397 tokens per probe throughout. SHA256 verified from local .sha256 file. Integrity Root from nuplan-2m-500q-summary-extract.
The ~92M avoided token figure is an estimate from record count times average token size. The per-probe bounded token figure (397) is measured.
RPS and throughput runs
| Run | Date | Records | Probes | Recall | Runtime | RPS | GPU W avg | SHA256 (prefix) |
|---|---|---|---|---|---|---|---|---|
| RPS Throughput 003 (50K) | 2026-07-05 | 50,000 | 50 / 50 | 100% | 16m 1s | 52.02 | 47.05W | 170025cd… |
| RPS Throughput 001 (100K) | 2026-07-04 | 100,000 | 100 / 100 | 100% | 32m 25s | 51.41 | 43.0W | 6868ccd3… |
| RPS RPS Profile 001 (500K) | 2026-07-04 | 500,000 | 500 / 500 | 100% | 4h 59m 9s | 27.86 | 34.70W | 6670c7c0… |
Low-power validation
Two independent 1M-record runs in low-power GPU configuration. Both achieved 100% verified retrieval recall at net incremental GPU draw typical of a workstation in idle state. Throughput did not degrade relative to standard config runs.
Low-power mode does not reduce throughput. The 47 RPS at 12.46W is faster than the 41.72 RPS at 16.51W in Ref 001. The power reduction is from GPU power state, not from reduced workload.
Cross-GPU validation
The retrieval architecture has been validated on A4000 Ada (RunPod), L40S (RunPod), H200 (RunPod), and H100 (RunPod) in addition to the primary RTX 4000 Ada 20GB workstation GPU. Cross-GPU runs confirm the architecture is not specific to a single GPU class.
GPU metric details: All GPU readings are from nvidia-smi.csv extracted per-run. Interval: 30 seconds. Available in all certified artifacts.
Claim boundary
These results prove that the TeleMemetry bounded retrieval system achieved 100% factual retrieval recall across up to 2,000,000 ingested nuPlan AV telemetry records, with bounded evidence packets of approximately 397 tokens per probe, in the documented run configurations.
This is a benchmark result on a specific dataset under specific conditions. It does not prove production safety certification, real-time AV control capability, or universal recall on arbitrary datasets.