Vercel Sandbox
5 harnesses
10 suites
Agent Benchmark Harness
Public leaderboard for Claude Code, Codex, OpenCode, Eve, and Mastra runs across coding, terminal, browser, OS, tool-use, finance, and skills benchmarks.
Completed Cells
0
10 failed, 15 infra failed
Average Score
n/a
completed cells
Average Runtime
n/a
completed cells
Recent Runs
9
last 50 stored runs
Top Scores
Highest normalized scores by suite, harness, and model.
Recent Runs
Latest run status and cell completion.
d9b76056-f754-49da-b8cc-f696e03463c8
15/15 cells
infra failed
6a6da63b-6647-465a-92da-1aa2b73d7347
10/150 cells
running
c03c0b80-f0e2-4700-bff4-4d1a3155b075
0/150 cells
running
f149a364-44b0-404a-ad0f-d7d8768c19ba
0/150 cells
running
acbe9348-af33-47df-aefc-70ae358848cf
0/150 cells
running
81d67420-2ece-4c6a-b4ae-2e1cc4eff5ee
0/150 cells
running
Harness Coverage
Average score by suite and harness, with sample counts in each cell.
Suite
claude-code
codex
opencode
eve
mastra
SWE-bench Verified
0.0%1
0.0%1
0.0%1
0.0%1
0.0%1
swe-bench-pro
n/a0
n/a0
n/a0
n/a0
n/a0
terminal-bench-2.1
n/a0
n/a0
n/a0
n/a0
n/a0
osworld-verified
n/a0
n/a0
n/a0
n/a0
n/a0
browsecomp
n/a0
n/a0
n/a0
n/a0
n/a0
mcp-atlas
n/a0
n/a0
n/a0
n/a0
n/a0
tau2-telecom
n/a0
n/a0
n/a0
n/a0
n/a0
finance-agent-v2
n/a0
n/a0
n/a0
n/a0
n/a0
vibe-code-bench-1.1
n/a0
n/a0
n/a0
n/a0
n/a0
skillsbench
n/a0
n/a0
n/a0
n/a0
n/a0
Leaderboard
Normalized result rows from completed and failed cells.
| Model | Harness | Suite | Score | Pass Rate | Cells | Avg Runtime |
|---|---|---|---|---|---|---|
opus-4.8-claude-code Opus 4.8 | claude-code | SWE-bench Verified | n/a | 0.0% | 2 | 302.5 ms |
opus-4.8-codex Opus 4.8 | codex | SWE-bench Verified | n/a | 0.0% | 2 | 113 ms |
opus-4.8-opencode Opus 4.8 | opencode | SWE-bench Verified | n/a | 0.0% | 2 | 176 ms |
opus-4.8-eve Opus 4.8 | eve | SWE-bench Verified | n/a | 0.0% | 2 | 142.5 ms |
opus-4.8-mastra Opus 4.8 | mastra | SWE-bench Verified | n/a | 0.0% | 2 | 110 ms |