A leaderboard of common benchmarks for each of our sandbox providers.
Powered byProvider | Score | Median | P95 | P99 | Success |
|---|---|---|---|---|---|
| Isorun | 99.4 | 0.06s | 0.07s | 0.07s | 100% |
| CreateOS | 98.2 | 0.17s | 0.18s | 0.18s | 100% |
| OpenComputer | 1.0 | 0.23s | 0.23s | 0.23s | 1% |
| Declaw | 96.4 | 0.31s | 0.42s | 0.46s | 100% |
| Archil | 95.2 | 0.45s | 0.53s | 0.54s | 100% |
| Arker | 94.1 | 0.56s | 0.64s | 0.65s | 100% |
| Beam | 93.2 | 0.60s | 0.80s | 0.80s | 100% |
| Vercel | 90.9 | 0.66s | 1.05s | 1.68s | 100% |
| Miosa | 92.6 | 0.70s | 0.79s | 0.82s | 100% |
| Tenki | 91.8 | 0.80s | 0.86s | 0.86s | 100% |
| Modal | 90.5 | 0.88s | 0.98s | 1.18s | 100% |
| Superserve | 88.1 | 1.00s | 1.44s | 1.54s | 100% |
| Freestyle | 87.7 | 1.11s | 1.40s | 1.43s | 100% |
| Runloop | 87.4 | 1.24s | 1.29s | 1.30s | 100% |
| E2B | 82.5 | 1.55s | 1.98s | 2.16s | 100% |
| Tensorlake | 79.7 | 1.97s | 2.10s | 2.12s | 100% |
| Sail | 73.7 | 2.49s | 2.84s | 2.88s | 100% |
| Cloud Run | 64.8 | 2.79s | 4.49s | 4.82s | 100% |
| Mosaic | 53.5 | 4.51s | 4.74s | 5.02s | 100% |
| Upstash | 40.4 | 4.82s | 7.50s | 7.95s | 100% |
| Cloudflare | 41.1 | 5.18s | 6.31s | 6.59s | 95% |
| givemeanode | 45.5 | 5.28s | 5.71s | 5.72s | 100% |
| Blaxel | 22.9 | 6.83s | 8.32s | 8.35s | 89% |
| CodeSandbox | 19.2 | 6.98s | 9.57s | 10.01s | 100% |
| Run Cloud | 12.6 | 7.91s | 19.02s | 23.25s | 100% |
| Sandbox0 | 0.0 | 12.66s | 12.75s | 12.75s | 5% |
| Daytona | 0.0 | 44.45s | 49.13s | 49.13s | 12% |
| Microsandbox | 0.0 | 61.93s | 110.65s | 115.78s | 94% |
Want to see a provider added?
Every benchmark measures Time to Interactive (TTI) — the elapsed time from calling compute.sandbox.create() to the first successful runCommand() inside the sandbox.
Each provider is tested with 100 iterations per run. Benchmarks run automatically via GitHub Actions on a recurring schedule. All results are committed to the public benchmarks repo.
Burst Test: All sandboxes are launched concurrently in a single burst.
The Composite Score is a weighted blend of timing metrics multiplied by the success rate. Each metric is scored against a fixed 10-second ceiling: 100 × (1 − value / 10,000ms), so a 200ms median scores 98 and anything ≥10s scores 0.
The weighted timing score is then multiplied by the success rate (0–1), so providers that fail frequently are penalized proportionally.
Have another question? Email us.
A sandbox is anywhere you can run code in isolation. It could be a VM, bare metal, a container, anywhere with compute resources.