Skip to content
GitHub

Product

Benchmarks Partners

Resources

Docs Blog

Burst TTI Benchmarks

Google Cloud Run logo Agents, sandboxes, & vibe-coded apps are powered by Cloud Run.

Burst TTI Benchmarks

A leaderboard of sandbox providers measured by time to interactive in a 100-sandbox burst.

Powered by Namespace · 4 vCPU / 16GB RAM · Northern Virginia, US
Last run: August 28, 2026

Performance Over Time

Composite Score

Detailed Metrics

Provider
Score
Median
P95
P99
Success
Isorun99.30.06s0.08s0.08s100%
CreateOS97.80.21s0.23s0.23s100%
Archil97.00.26s0.35s0.37s100%
OpenComputer83.60.27s0.29s0.30s86%
Arker96.70.31s0.35s0.36s100%
Daytona87.50.33s0.48s0.49s91%
Declaw94.80.48s0.57s0.58s100%
Superserve93.70.49s0.81s0.87s100%
Blaxel94.50.52s0.59s0.63s100%
Vercel93.80.54s0.70s0.83s100%
Cloud Run91.00.68s1.17s1.29s100%
Miosa92.40.70s0.85s0.86s100%
Modal91.10.85s0.95s0.96s100%
Mosaic83.30.97s1.48s1.84s95%
Runloop89.11.06s1.13s1.13s100%
E2B88.21.09s1.31s1.37s100%
Beam84.51.34s1.65s1.65s99%
Tensorlake81.81.40s2.43s2.43s100%
Tenki81.31.67s2.16s2.17s100%
Sail72.02.68s2.96s2.99s100%
Cloudflare40.65.26s6.71s7.41s100%
Upstash36.55.27s7.94s8.01s100%
CodeSandbox18.26.96s10.04s10.64s100%
Run Cloud0.011.81s22.60s26.35s100%
Sandbox00.027.38s27.49s27.52s24%
Hopx0.00.00s0.00s0.00s0%
Lightning AI0.00.00s0.00s0.00s0%
Microsandbox0.00.00s0.00s0.00s0%
Northflank0.00.00s0.00s0.00s0%

Want to see a provider added?

Let us know on X

Methodology

What We Measure

Every benchmark measures Time to Interactive (TTI) — the elapsed time from calling compute.sandbox.create() to the first successful runCommand() inside the sandbox.

Each provider is tested with 100 iterations per run. Benchmarks run automatically via GitHub Actions on a recurring schedule. All results are committed to the public benchmarks repo.

Burst Test: All sandboxes are launched concurrently in a single burst.

How We Score

The Composite Score is a weighted blend of timing metrics multiplied by the success rate. Each metric is scored against a fixed 10-second ceiling: 100 × (1 − value / 10,000ms), so a 200ms median scores 98 and anything ≥10s scores 0.

The weighted timing score is then multiplied by the success rate (0–1), so providers that fail frequently are penalized proportionally.

  • Median: 60% — primary signal for typical experience
  • P95: 25% — tail latency / consistency
  • P99: 15% — extreme tail latency

Sandbox Benchmarks FAQs

Have another question? Email us.

A sandbox is anywhere you can run code in isolation. It could be a VM, bare metal, a container, anywhere with compute resources.
PartnersLatitude