Sandbox Benchmarks - Provider Leaderboard
Sandbox Benchmarks
A leaderboard of common benchmarks for each of our sandbox providers.
Powered byPerformance Over Time
Composite Score
Detailed Metrics
Provider | Score | Median | P95 | P99 | Success |
|---|---|---|---|---|---|
| Isorun | 97.1 | 0.29s | 0.29s | 0.29s | 100% |
| Northflank | 96.9 | 0.29s | 0.34s | 0.35s | 100% |
| Lightning AI | 96.2 | 0.33s | 0.45s | 0.47s | 100% |
| Beam | 96.4 | 0.33s | 0.40s | 0.41s | 100% |
| Blaxel | 96.3 | 0.34s | 0.41s | 0.41s | 100% |
| CreateOS | 95.8 | 0.39s | 0.46s | 0.47s | 100% |
| Declaw | 94.4 | 0.44s | 0.71s | 0.79s | 100% |
| Vercel | 93.6 | 0.53s | 0.77s | 0.83s | 100% |
| Cloud Run | 91.8 | 0.63s | 1.08s | 1.16s | 100% |
| Modal | 92.8 | 0.70s | 0.74s | 0.76s | 100% |
| Archil | 92.1 | 0.70s | 0.91s | 0.92s | 100% |
| Runloop | 91.7 | 0.80s | 0.87s | 0.87s | 100% |
| Tensorlake | 91.3 | 0.85s | 0.90s | 0.91s | 100% |
| E2B | 90.0 | 0.94s | 1.10s | 1.12s | 100% |
| Superserve | 88.8 | 0.94s | 1.39s | 1.43s | 100% |
| Tenki | 85.0 | 1.42s | 1.60s | 1.64s | 100% |
| Daytona | 49.2 | 1.81s | 13.70s | 14.62s | 100% |
| Upstash | 72.4 | 2.10s | 3.74s | 3.76s | 100% |
| OpenComputer | 70.7 | 2.59s | 3.43s | 3.45s | 100% |
| Cloudflare | 49.4 | 4.53s | 5.77s | 6.03s | 100% |
| CodeSandbox | 13.2 | 7.80s | 10.57s | 10.93s | 100% |
| Sandbox0 | 0.0 | 13.04s | 23.37s | 23.91s | 59% |
Want to see a provider added?
Methodology
What We Measure
Every benchmark measures Time to Interactive (TTI) — the elapsed time from calling compute.sandbox.create() to the first successful runCommand() inside the sandbox.
Each provider is tested with 100 iterations per run. Benchmarks run automatically via GitHub Actions on a recurring schedule. All results are committed to the public benchmarks repo.
Sequential Test: Sandboxes are launched one at a time, waiting for each to become interactive before starting the next.
Staggered Test: Sandboxes are launched with 200ms delays between each.
Burst Test: All sandboxes are launched concurrently in a single burst.
How We Score
The Composite Score is a weighted blend of timing metrics multiplied by the success rate. Each metric is scored against a fixed 10-second ceiling: 100 × (1 − value / 10,000ms), so a 200ms median scores 98 and anything ≥10s scores 0.
The weighted timing score is then multiplied by the success rate (0–1), so providers that fail frequently are penalized proportionally.
- • Median: 60% — primary signal for typical experience
- • P95: 25% — tail latency / consistency
- • P99: 15% — extreme tail latency
Sandbox Benchmarks FAQs
Have another question? Email us.