Skip to content
GitHub
Compute SDK

Product

Benchmarks Partners

Resources

Docs Blog

Sandbox Benchmarks - Provider Leaderboard

Google Cloud Run logo Agents, sandboxes, & vibe-coded apps are powered by Cloud Run.

Sandbox Benchmarks

A leaderboard of common benchmarks for each of our sandbox providers.

Powered by Namespace · 4 vCPU / 16GB RAM · Northern Virginia, US
Last run: July 31, 2026

Performance Over Time

Composite Score

Detailed Metrics

Provider
Score
Median
P95
P99
Success
Isorun97.10.29s0.29s0.29s100%
Northflank96.90.29s0.34s0.35s100%
Lightning AI96.20.33s0.45s0.47s100%
Beam96.40.33s0.40s0.41s100%
Blaxel96.30.34s0.41s0.41s100%
CreateOS95.80.39s0.46s0.47s100%
Declaw94.40.44s0.71s0.79s100%
Vercel93.60.53s0.77s0.83s100%
Cloud Run91.80.63s1.08s1.16s100%
Modal92.80.70s0.74s0.76s100%
Archil92.10.70s0.91s0.92s100%
Runloop91.70.80s0.87s0.87s100%
Tensorlake91.30.85s0.90s0.91s100%
E2B90.00.94s1.10s1.12s100%
Superserve88.80.94s1.39s1.43s100%
Tenki85.01.42s1.60s1.64s100%
Daytona49.21.81s13.70s14.62s100%
Upstash72.42.10s3.74s3.76s100%
OpenComputer70.72.59s3.43s3.45s100%
Cloudflare49.44.53s5.77s6.03s100%
CodeSandbox13.27.80s10.57s10.93s100%
Sandbox00.013.04s23.37s23.91s59%

Want to see a provider added?

Let us know on X

Methodology

What We Measure

Every benchmark measures Time to Interactive (TTI) — the elapsed time from calling compute.sandbox.create() to the first successful runCommand() inside the sandbox.

Each provider is tested with 100 iterations per run. Benchmarks run automatically via GitHub Actions on a recurring schedule. All results are committed to the public benchmarks repo.

Sequential Test: Sandboxes are launched one at a time, waiting for each to become interactive before starting the next.

Staggered Test: Sandboxes are launched with 200ms delays between each.

Burst Test: All sandboxes are launched concurrently in a single burst.

How We Score

The Composite Score is a weighted blend of timing metrics multiplied by the success rate. Each metric is scored against a fixed 10-second ceiling: 100 × (1 − value / 10,000ms), so a 200ms median scores 98 and anything ≥10s scores 0.

The weighted timing score is then multiplied by the success rate (0–1), so providers that fail frequently are penalized proportionally.

  • Median: 60% — primary signal for typical experience
  • P95: 25% — tail latency / consistency
  • P99: 15% — extreme tail latency

Sandbox Benchmarks FAQs

Have another question? Email us.

A sandbox is anywhere you can run code in isolation. It could be a VM, bare metal, a container, anywhere with compute resources.
PartnersLatitude