2,000,000
Max records · completed run
Endurance 001 · continuous · 33h 38m
100%
Recall · all certified runs
Measured · 0 integrity failures
~397
Tokens per probe
Measured · Ref 002 config · fixed regardless of archive depth
52.02
Peak RPS
Throughput 003 · 50K · 47.05W avg GPU

Dataset

nuPlan Mini AV dataset

nuPlan is a large-scale AV planning dataset from Motional / nuTonomy. The Mini subset used in TeleMemetry benchmarks contains real nuPlan telemetry replayed from stored files as a streaming feed. It is not live vehicle data, not an NVIDIA DRIVE integration, and not a production safety validation.

Data lane: REAL NUPLAN SUMMARY. No raw telemetry frames are included in public artifacts. Numbers in result cards come from the run summary JSON and were not adjusted or normalized after the fact.

Dataset not included in public artifacts. Independent reviewers may request access to the dataset configuration and replay harness at AIFSOfficial@proton.me.

Canonical reference runs

Run Date Records Probes Recall Tok/Q Runtime RPS GPU W avg VRAM peak SHA256 (prefix) Integrity Root (prefix)
REF Ref 001 2026-06-29 1,000,000 100 / 100 100% 397 6h 39m 31s 41.72 16.51W 4.73 GB 5a06e018… 8ee1a1bf…
REF Ref 002 2026-06-30 1,000,000 500 / 500 100% 397 11h 29m 55s 24.16 25.32W 4.895 GB 37735071… bd2e0eec…
REF Ref 003 2026-07-03 1,000,000 500 / 500 100% 397 11h 55m 8s 23.31 24.91W n/a 2e8c0f3e… n/a

Ref 001 notes

Initial 1M run. 94/100 answer-layer citation accuracy (retrieval 100/100). The 6 misses were at the answer/citation layer, not the retrieval layer. SHA256, Integrity Root, and GPU CSV verified locally.

Ref 002 notes

1M records, 500/500 retrieval recall. Second artifact also produced: gatekeeper-av-nuplan-1m-500q-rtx4000ada20gb-qwen25-7b-proofstress-20260630-131945.tgz (SHA256: c24bd4834f87843e730691cb3ff1841bef1bb0e605cdbf56fc2718bf7273999d). GPU CSV and Integrity Root verified locally.

Endurance run

Run Date Records Probes Recall Runtime RPS Replay tokens avoided (est.) SHA256 (prefix) Integrity Root (prefix)
END Endurance 001 2026-07-01 2,000,000 500 / 500 100% 33h 38m 16.51 ~92,000,000 4e6db889… ec0cbf43…

Endurance 001 notes

2M-record continuous run. Largest completed nuPlan run. At a hypothetical full-replay baseline of ~100 tokens/record, the 2M archive would require ~200M context tokens per probe. TeleMemetry used ~397 tokens per probe throughout. SHA256 verified from local .sha256 file. Integrity Root from nuplan-2m-500q-summary-extract.

The ~92M avoided token figure is an estimate from record count times average token size. The per-probe bounded token figure (397) is measured.

RPS and throughput runs

Run Date Records Probes Recall Runtime RPS GPU W avg SHA256 (prefix)
RPS Throughput 003 (50K) 2026-07-05 50,000 50 / 50 100% 16m 1s 52.02 47.05W 170025cd…
RPS Throughput 001 (100K) 2026-07-04 100,000 100 / 100 100% 32m 25s 51.41 43.0W 6868ccd3…
RPS RPS Profile 001 (500K) 2026-07-04 500,000 500 / 500 100% 4h 59m 9s 27.86 34.70W 6670c7c0…

Low-power validation

1M records at 12.46W avg GPU power

Two independent 1M-record runs in low-power GPU configuration. Both achieved 100% verified retrieval recall at net incremental GPU draw typical of a workstation in idle state. Throughput did not degrade relative to standard config runs.

12.53W
Low-Power Watch 001
2026-07-06 · 46.01 RPS · 50/50 recall
12.46W
Low-Power Watch 002
2026-07-07 · 47.05 RPS · 50/50 recall · 4.78 GB VRAM
46.01
RPS · Watch 001
vs 41.72 Ref 001 at 16.51W
47.05
RPS · Watch 002
Reproducibility confirmed

Low-power mode does not reduce throughput. The 47 RPS at 12.46W is faster than the 41.72 RPS at 16.51W in Ref 001. The power reduction is from GPU power state, not from reduced workload.

Cross-GPU validation

Other GPU classes tested

The retrieval architecture has been validated on A4000 Ada (RunPod), L40S (RunPod), H200 (RunPod), and H100 (RunPod) in addition to the primary RTX 4000 Ada 20GB workstation GPU. Cross-GPU runs confirm the architecture is not specific to a single GPU class.

GPU metric details: All GPU readings are from nvidia-smi.csv extracted per-run. Interval: 30 seconds. Available in all certified artifacts.

Claim boundary

What these results prove

These results prove that the TeleMemetry bounded retrieval system achieved 100% factual retrieval recall across up to 2,000,000 ingested nuPlan AV telemetry records, with bounded evidence packets of approximately 397 tokens per probe, in the documented run configurations.

This is a benchmark result on a specific dataset under specific conditions. It does not prove production safety certification, real-time AV control capability, or universal recall on arbitrary datasets.