Benchmark FAQ

A compact guide to TeleMemetry benchmark categories, evidence packages, and claim boundaries. Detailed domain pages remain available, but this page is the clean overview.

100,000 / 100,000
Bit-Perfect AI Turns
Exact operational recall over captured Isaac Lab telemetry.
3,548x
Replay Reduction
Fewer active tokens versus full replay estimate on the Isaac flagship.
278
Average Active Context
Tokens per verified AI turn on the Isaac flagship.
5
Evidence Lanes
Isaac, nuPlan, satellite, live CelesTrak, and alignment context.
T

TeleMemetry Benchmark FAQ

What a verified AI turn means · what is measured · what is not claimed

Core

Primary Unit

Verified AI Turn

Definition

Exact operational recall request + AI response + deterministic verification.

Public Standard

Bit-perfect, field-for-field, no partial credit.

Boundary

TeleMemetry benchmarks operational memory and verified recall. It does not claim model superiority, universal reasoning, or physical-world task success from these benchmark rows.

Open registry
I

Isaac Lab Robotics Telemetry

Captured simulation archive · bit-perfect AI turns · not a physics-throughput benchmark

Flagship

Headline

100,000 / 100,000 Bit-Perfect AI Turns

Archive

Captured Isaac Lab telemetry, probed later.

Efficiency

3,548x replay reduction, 278 active tokens / turn.

Boundary

This is a telemetry recall benchmark over Isaac-derived data. It does not measure Isaac Lab physics stepping speed, robotics reasoning, control quality, policy success, or physical task success.

Registry evidence
N

nuPlan AV Telemetry

Automotive telemetry domain · million-record scale · artifact-backed runs

Certified

Max Scale

2,000,000 records

Canonical Recall

100% verified recall across certified reference runs.

Active Context

~397 tokens / request in reference configs.

Boundary

nuPlan uses stored AV telemetry replay. It is not live vehicle data, not an NVIDIA DRIVE integration, and not a production safety validation.

S

Satellite Recall and Live CelesTrak

Certified satellite recall runs kept separate from transient live demonstrations

Domain

Certified Lane

Packaged satellite recall runs with SHA256 receipts.

Live Lane

CelesTrak session snapshots, transient by design.

Why Split

Certified artifacts and live demos carry different proof strength.

Boundary

Live CelesTrak rows are demonstrations of operational behavior during uptime. They are not sealed benchmark artifacts and reset with the service session.

Registry evidence Live demo
L

LongMemEval

Conversational memory validation · secondary domain · not the headline product claim

Secondary

Purpose

Tests whether bounded retrieval generalizes beyond deterministic telemetry.

Dataset

LongMemEval V1 conversational memory benchmark.

Status

Secondary evidence lane.

Boundary

LongMemEval is useful for architectural generalization, but it should not be presented as the primary TeleMemetry telemetry proof.

Detailed page
M

MLPerf Alignment

Inference benchmarks and operational memory benchmarks measure different layers

Context

MLPerf

Model/hardware inference throughput and latency.

TeleMemetry

Verified operational memory and bounded active context.

Relationship

Complementary, not competing.

Boundary

TeleMemetry has not submitted to MLPerf. The comparison explains workload placement, not an official MLCommons result.

Detailed page

Claim Boundary

TeleMemetry evidence packages support verified operational memory claims within each documented run configuration. They do not prove model reasoning quality, universal memory, robotics control performance, physical task success, or production safety readiness by themselves.