Dax Benchmark

A real-world workload benchmark inspired by this tweet from dax to test potential sandbox providers for OpenCode.

Per-Phase Duration Breakdown

Median duration per phase of the clone + install + typecheck cycle.

Detailed Metrics

Provider
Phases
Success
Total (med)
Prepare
Bun DL
Bun Unpack
Clone
Install
Typecheck
Isorun7/7100%33.54s1.41s0.22s0.43s1.11s8.06s19.57s
Blaxel7/7100%43.75s1.24s0.44s0.48s1.51s9.47s23.46s
Namespace7/7100%45.89s3.22s0.25s0.45s1.10s10.72s23.65s
CreateOS7/7100%48.13s2.07s0.59s0.45s2.53s12.53s25.87s
Upstash7/7100%49.41s3.68s0.26s0.62s1.66s9.90s24.04s
Mosaic7/7100%59.71s1.20s0.31s0.63s2.31s13.88s30.38s
Arker7/7100%60.28s8.91s0.36s0.64s2.00s11.74s32.60s
Vercel7/7100%66.84s11.13s0.25s0.69s2.08s11.31s35.33s
Daytona7/7100%73.38s5.33s0.36s0.60s1.70s14.93s37.60s
Tenki7/7100%79.91s4.61s0.56s0.72s3.16s20.98s39.39s
Tensorlake7/7100%85.42s17.15s0.43s1.06s3.34s17.62s39.79s
Modal7/7100%94.26s8.57s0.38s0.70s2.36s14.03s61.26s
Sail7/7100%94.31s4.35s1.02s0.89s3.90s18.95s55.74s
Superserve7/7100%109.94s12.18s0.40s1.07s6.22s19.55s67.31s
Runloop7/7100%120.89s16.89s0.41s0.80s3.72s39.94s53.77s
Sandbox07/7100%171.98s29.31s1.68s0.86s8.81s73.22s36.88s
Declaw7/7100%277.34s5.65s0.71s0.78s4.95s48.25s202.77s
E2B6/70%Failed10.90s0.83s0.78s2.92s
Microsandbox4/70%Failed2.54s0.31s
OpenComputer3/70%Failed7.34s
Cloud Run2/70%Failed
Archil0/70%Failed
Beam0/70%Failed
Cloudflare0/70%Failed
CodeSandbox0/70%Failed
Freestyle0/70%Failed
givemeanode0/70%Failed
Hopx0/70%Failed
Lightning AI0/70%Failed
Miosa0/70%Failed
Northflank0/70%Failed
Run Cloud0/70%Failed

Performance Over Time

Want to see a provider added?

Let us know on X

Methodology

What We Measure

Each run executes a cold clone + install + typecheck cycle of the opencode repository inside a fresh sandbox: install system packages and Node.js, download and unpack Bun, shallow-clone opencode at a pinned commit, run bun install, then bun typecheck.

Each provider runs 1 iteration per scheduled benchmark run. Benchmarks run via GitHub Actions and results are committed to the public benchmarks repo.

How We Score

The Composite Score is the median number of the 7 benchmark phases (prepare, cache clear, bun download, bun unpack, clone, install, typecheck) a provider completed before failing, if any — shown as a fraction like 7/7.

The other metrics (Total, Prepare, Clone, Install, Typecheck) are real-world phase durations in seconds — providers are ranked by median duration (lower is better) when one of those is selected. The benchmark requires curl, root/sudo, and a package manager (apt, dnf, or apk) inside the sandbox — providers without those capabilities fail early, which shows up as a low phase count.

PartnersLatitude