Engine research began mid-2025 · Formal artifact-backed runs: June 2026 · Hardware: RTX 4000 Ada 20GB · H100 · H200 · L40S
| Run | Recall | Probes | Avg Latency | Max Latency | Board Power | Max Temp | Runtime | Evidence | Type | SHA256 |
|---|---|---|---|---|---|---|---|---|---|---|
| Overnight Endurance RTX 4000 Ada · qwen2.5:7b · 1 sec cadence · no pauses |
100% | 18,821 / 18,821 | 915.9 ms | 5,009 ms | 74.83 W | 62 °C | 6h 00m 10s | ★ Canonical SHA256 ✓ | ENDURANCE | 44b5c488…772d |
| Max Sustained Recall RTX 4000 Ada · qwen2.5:7b · 0 ms delay |
100% | 2,987 / 2,987 | 915.2 ms | 1,095 ms | 74.9 W | 61 °C | 1 hour | SHA256 ✓ | ENDURANCE | 99f58352…c61f |
| Balanced Operational Recall RTX 4000 Ada · qwen2.5:7b · 5 sec cadence |
100% | 699 / 699 | 933.7 ms | 5,030 ms | 60.15 W | 54 °C | 1 hour | SHA256 ✓ | ENDURANCE | 2db82b96…927 |
| Total Punishment Run RTX 4000 Ada · qwen2.5:7b · 0 ms delay · 15 min |
100% | 795 / 795 | 919 ms | 5,022 ms | 75.46 W | 61 °C | 15 min | SHA256 ✓ | STRESS | e66ee4dd…c2a |
This index lists satellite telemetry benchmark artifacts and measured results. It does not certify general reasoning ability, production control safety, or live-feed performance outside the packaged artifacts.
Satellite fields probed: satellite_id, timestamp, latitude, longitude, altitude_km, speed_km_s, heading_deg. Exact byte-for-byte return required for a PASS.
Board-power maxima are listed as audit telemetry, not inference-power headline claims. ~0.017 Wh is calculated from ~52.8 W estimated inference load above loaded-idle baseline across 18,821 verified recalls.
Click any row to open the full run record including SHA256 and copy-to-clipboard.
| Run | Full | Summary | Purpose | Records | GPU | AI Model | Recall | Active Tokens | Avg GPU W | Avg RPS | Replay Avoided | Runtime | Efficiency | Evidence | Type | Date | Issues | SHA256 |
|---|
This index lists benchmark artifacts and measured results. It does not claim general AI reasoning improvement, model superiority, or universal dataset performance.
nuPlan dataset: nuPlan Mini telemetry replayed from stored files as a streaming feed. Not live vehicle data, not an NVIDIA DRIVE integration, not production safety validation.
Planted recall / debate dataset: Synthetic conversation sequences with inserted recall targets. Gatekeeper recall measures whether the bounded retrieval system surfaced the correct planted record.
Replay tokens avoided: Estimated tokens that full transcript replay would have consumed at each challenge point. Computed from run summaries. Labeled as estimated unless a concurrent replay run produced the baseline. The ref-001 and ref-002 values (both 46M) come from the same estimation methodology and have not been independently recomputed from artifact energy windows. The endurance-001 value (92M) reflects approximately 2x the reference run figure at 2x the scale.
Canonical runs (marked ★) are the primary evidence artifacts. Other runs document engineering progression.
| Snapshot | Type | Recall | Probes | Events Ingested | Satellites | Avg Context | Replay Avoided | Status |
|---|---|---|---|---|---|---|---|---|
| Live Proof Banner Snapshot CelesTrak satellite orbital data · 678 AI History probes |
LIVE DEMO | 99.9% | 677 / 678 | 107,231 | 157 | 583 tokens | ~2.4B | Transient · not sealed |
| Live Audit Modal Snapshot Public service session · continuously updated |
LIVE DEMO | 100% | 1,329 / 1,329 | · | · | · | · | Transient · not sealed |
CelesTrak values are real public-service session snapshots. They are transient and reset when the service restarts. Do not cite these counts as certified deterministic benchmark results.
Certified sealed artifact results are in the Certified nuPlan Runs and RunPod Local Satellite Recall Runs panels.
Click any row to open the snapshot record.
| Run | Dataset | Questions | Retrieval Recall | Answer F1 | Active Tokens | Evidence | Type | Date |
|---|---|---|---|---|---|---|---|---|
| No certified runs yet | ||||||||
| Run | Dataset | Questions | Retrieval Recall | Answer F1 | Active Tokens | Evidence | Type | Date |
|---|---|---|---|---|---|---|---|---|
| No certified runs yet | ||||||||
| Run | Benchmark | Metric | Score | Active Tokens | Evidence | Type | Date |
|---|---|---|---|---|---|---|---|
| No certified runs yet | |||||||