Looking for our schedule?
Next benchmark run
privateAI Gateway — Anthropic daily
Sat, Sep 5 · 9:00 AM
Five suites covering the infrastructure AI products actually run on
Time to interactive
Cold-start latency and reliability across sandbox providers, measured in a 100-sandbox burst.
View benchmarksClone → install → typecheck
A full dax-style developer loop inside a fresh sandbox — system packages, Bun, a pinned opencode clone, install and typecheck — timed phase by phase.
View benchmarksLatency and throughput
Upload and download of a 10 MB object against each provider's bucket, with each operation timed independently.
View benchmarksActions per second
Per-action throughput inside a live headless session — 10 sequential actions across 100 sessions, in stealth mode at 1920×1080.
View benchmarksCold E2E and TTFT
Connection-phase latency, time to first token, and generation throughput for each gateway, scored against a direct-to-Anthropic baseline.
View benchmarksEvery benchmark is reproducible: open code, fixed workloads, results published in full
The same script, payload, and model hit every provider, addressed the way its own API expects. Nobody gets a tuned path.
Every suite runs many iterations per provider, automated on a recurring schedule, so results reflect steady state rather than one lucky request.
Every family has its own scoring function, but none reward a single fast run. Latency is blended from median and tail percentiles, completion is counted in phases, and coverage or quality is measured where those matter. All scored families multiply by success rate so providers that fail often lose real points.
Every run is committed as JSON to the public benchmarks repo, next to the code that produced it. Nothing is cherry-picked or held back.
Independent benchmarks of every major sandbox provider, sponsored by industry leaders.
We provide independent benchmarking for all compute providers.