Google Cloud Run logoAgents, sandboxes, & vibe-coded apps are powered by Cloud Run.

Sandbox Benchmarks

A leaderboard of common benchmarks for each of our sandbox providers.

Powered byNamespace· 4 vCPU / 16GB RAM· Northern Virginia, US
Last run: September 4, 2026

Performance Over Time

Composite Score

Detailed Metrics

Provider
Score
Median
P95
P99
Success
Isorun99.40.06s0.07s0.07s100%
CreateOS98.20.17s0.18s0.18s100%
OpenComputer1.00.23s0.23s0.23s1%
Declaw96.40.31s0.42s0.46s100%
Archil95.20.45s0.53s0.54s100%
Arker94.10.56s0.64s0.65s100%
Beam93.20.60s0.80s0.80s100%
Vercel90.90.66s1.05s1.68s100%
Miosa92.60.70s0.79s0.82s100%
Tenki91.80.80s0.86s0.86s100%
Modal90.50.88s0.98s1.18s100%
Superserve88.11.00s1.44s1.54s100%
Freestyle87.71.11s1.40s1.43s100%
Runloop87.41.24s1.29s1.30s100%
E2B82.51.55s1.98s2.16s100%
Tensorlake79.71.97s2.10s2.12s100%
Sail73.72.49s2.84s2.88s100%
Cloud Run64.82.79s4.49s4.82s100%
Mosaic53.54.51s4.74s5.02s100%
Upstash40.44.82s7.50s7.95s100%
Cloudflare41.15.18s6.31s6.59s95%
givemeanode45.55.28s5.71s5.72s100%
Blaxel22.96.83s8.32s8.35s89%
CodeSandbox19.26.98s9.57s10.01s100%
Run Cloud12.67.91s19.02s23.25s100%
Sandbox00.012.66s12.75s12.75s5%
Daytona0.044.45s49.13s49.13s12%
Microsandbox0.061.93s110.65s115.78s94%

Want to see a provider added?

Let us know on X

Methodology

What We Measure

Every benchmark measures Time to Interactive (TTI) — the elapsed time from calling compute.sandbox.create() to the first successful runCommand() inside the sandbox.

Each provider is tested with 100 iterations per run. Benchmarks run automatically via GitHub Actions on a recurring schedule. All results are committed to the public benchmarks repo.

Burst Test: All sandboxes are launched concurrently in a single burst.

How We Score

The Composite Score is a weighted blend of timing metrics multiplied by the success rate. Each metric is scored against a fixed 10-second ceiling: 100 × (1 − value / 10,000ms), so a 200ms median scores 98 and anything ≥10s scores 0.

The weighted timing score is then multiplied by the success rate (0–1), so providers that fail frequently are penalized proportionally.

  • Median: 60% — primary signal for typical experience
  • P95: 25% — tail latency / consistency
  • P99: 15% — extreme tail latency

Sandbox Benchmarks FAQs

Have another question? Email us.

A sandbox is anywhere you can run code in isolation. It could be a VM, bare metal, a container, anywhere with compute resources.

PartnersLatitude