Burst TTI Benchmarks
Burst TTI Benchmarks
A leaderboard of sandbox providers measured by time to interactive in a 100-sandbox burst.
Powered byProvider Leaderboard
Performance Over Time
Composite Score
Detailed Metrics
Provider | Score | Median | P95 | P99 | Success |
|---|---|---|---|---|---|
| Isorun | 99.3 | 0.06s | 0.08s | 0.08s | 100% |
| CreateOS | 97.8 | 0.21s | 0.23s | 0.23s | 100% |
| Archil | 97.0 | 0.26s | 0.35s | 0.37s | 100% |
| OpenComputer | 83.6 | 0.27s | 0.29s | 0.30s | 86% |
| Arker | 96.7 | 0.31s | 0.35s | 0.36s | 100% |
| Daytona | 87.5 | 0.33s | 0.48s | 0.49s | 91% |
| Declaw | 94.8 | 0.48s | 0.57s | 0.58s | 100% |
| Superserve | 93.7 | 0.49s | 0.81s | 0.87s | 100% |
| Blaxel | 94.5 | 0.52s | 0.59s | 0.63s | 100% |
| Vercel | 93.8 | 0.54s | 0.70s | 0.83s | 100% |
| Cloud Run | 91.0 | 0.68s | 1.17s | 1.29s | 100% |
| Miosa | 92.4 | 0.70s | 0.85s | 0.86s | 100% |
| Modal | 91.1 | 0.85s | 0.95s | 0.96s | 100% |
| Mosaic | 83.3 | 0.97s | 1.48s | 1.84s | 95% |
| Runloop | 89.1 | 1.06s | 1.13s | 1.13s | 100% |
| E2B | 88.2 | 1.09s | 1.31s | 1.37s | 100% |
| Beam | 84.5 | 1.34s | 1.65s | 1.65s | 99% |
| Tensorlake | 81.8 | 1.40s | 2.43s | 2.43s | 100% |
| Tenki | 81.3 | 1.67s | 2.16s | 2.17s | 100% |
| Sail | 72.0 | 2.68s | 2.96s | 2.99s | 100% |
| Cloudflare | 40.6 | 5.26s | 6.71s | 7.41s | 100% |
| Upstash | 36.5 | 5.27s | 7.94s | 8.01s | 100% |
| CodeSandbox | 18.2 | 6.96s | 10.04s | 10.64s | 100% |
| Run Cloud | 0.0 | 11.81s | 22.60s | 26.35s | 100% |
| Sandbox0 | 0.0 | 27.38s | 27.49s | 27.52s | 24% |
| Hopx | 0.0 | 0.00s | 0.00s | 0.00s | 0% |
| Lightning AI | 0.0 | 0.00s | 0.00s | 0.00s | 0% |
| Microsandbox | 0.0 | 0.00s | 0.00s | 0.00s | 0% |
| Northflank | 0.0 | 0.00s | 0.00s | 0.00s | 0% |
Want to see a provider added?
Methodology
What We Measure
Every benchmark measures Time to Interactive (TTI) — the elapsed time from calling compute.sandbox.create() to the first successful runCommand() inside the sandbox.
Each provider is tested with 100 iterations per run. Benchmarks run automatically via GitHub Actions on a recurring schedule. All results are committed to the public benchmarks repo.
Burst Test: All sandboxes are launched concurrently in a single burst.
How We Score
The Composite Score is a weighted blend of timing metrics multiplied by the success rate. Each metric is scored against a fixed 10-second ceiling: 100 × (1 − value / 10,000ms), so a 200ms median scores 98 and anything ≥10s scores 0.
The weighted timing score is then multiplied by the success rate (0–1), so providers that fail frequently are penalized proportionally.
- • Median: 60% — primary signal for typical experience
- • P95: 25% — tail latency / consistency
- • P99: 15% — extreme tail latency
Sandbox Benchmarks FAQs
Have another question? Email us.