Create secure sandboxes anywhere
Independent benchmarks for compute
The best way to evaluate your compute provider performance.
Median TTI (burst) over time
Sandboxes
What We Benchmark
Five suites covering the infrastructure AI products actually run on
Time to interactive
Sandboxes
Cold-start latency and reliability across sandbox providers, measured sequentially, in a 100-sandbox burst, and under staggered load.
View benchmarksClone → install → typecheck
Real-world workload
A full dax-style developer loop inside a fresh sandbox — system packages, Bun, a pinned opencode clone, install and typecheck — timed phase by phase.
View benchmarksLatency and throughput
Storage
Upload and download of a 10 MB object against each provider's bucket, with each operation timed independently.
View benchmarksActions per second
Browsers
Per-action throughput inside a live headless session — 50 sequential actions across 10 sessions, in stealth mode at 1920×1080.
View benchmarksCold E2E and TTFT
AI gateways
Connection-phase latency, time to first token, and generation throughput for each gateway, scored against a direct-to-Anthropic baseline.
View benchmarksHow the Numbers Are Made
Every benchmark is reproducible: open code, fixed workloads, results published in full
One workload, every provider
The same script, payload, and model hit every provider, addressed the way its own API expects. Nobody gets a tuned path.
Repeated on a schedule
Every suite runs many iterations per provider, automated in GitHub Actions on a recurring schedule, so results reflect steady state rather than one lucky request.
Scored on the tail, not the best case
Each composite score blends median and tail latency against a fixed ceiling, then multiplies by success rate — providers that fail often lose real points.
Raw results published
Every run is committed as JSON to the public benchmarks repo, next to the code that produced it. Nothing is cherry-picked or held back.
Trusted by the best
Independent benchmarks of every major sandbox provider, sponsored by industry leaders.
Start Benchmarking
We provide independent benchmarking for all compute providers.