Vercel Sandbox
5 harnesses
10 suites

Agent Benchmark Harness

Public leaderboard for Claude Code, Codex, OpenCode, Eve, and Mastra runs across coding, terminal, browser, OS, tool-use, finance, and skills benchmarks.

Completed Cells
0
10 failed, 15 infra failed
Average Score
n/a
completed cells
Average Runtime
n/a
completed cells
Recent Runs
9
last 50 stored runs
Top Scores
Highest normalized scores by suite, harness, and model.
Recent Runs
Latest run status and cell completion.
d9b76056-f754-49da-b8cc-f696e03463c8
15/15 cells
infra failed
6a6da63b-6647-465a-92da-1aa2b73d7347
10/150 cells
running
c03c0b80-f0e2-4700-bff4-4d1a3155b075
0/150 cells
running
f149a364-44b0-404a-ad0f-d7d8768c19ba
0/150 cells
running
acbe9348-af33-47df-aefc-70ae358848cf
0/150 cells
running
81d67420-2ece-4c6a-b4ae-2e1cc4eff5ee
0/150 cells
running
Harness Coverage
Average score by suite and harness, with sample counts in each cell.
Suite
claude-code
codex
opencode
eve
mastra
SWE-bench Verified
0.0%1
0.0%1
0.0%1
0.0%1
0.0%1
swe-bench-pro
n/a0
n/a0
n/a0
n/a0
n/a0
terminal-bench-2.1
n/a0
n/a0
n/a0
n/a0
n/a0
osworld-verified
n/a0
n/a0
n/a0
n/a0
n/a0
browsecomp
n/a0
n/a0
n/a0
n/a0
n/a0
mcp-atlas
n/a0
n/a0
n/a0
n/a0
n/a0
tau2-telecom
n/a0
n/a0
n/a0
n/a0
n/a0
finance-agent-v2
n/a0
n/a0
n/a0
n/a0
n/a0
vibe-code-bench-1.1
n/a0
n/a0
n/a0
n/a0
n/a0
skillsbench
n/a0
n/a0
n/a0
n/a0
n/a0
Leaderboard
Normalized result rows from completed and failed cells.
ModelHarnessSuiteScorePass RateCellsAvg Runtime
opus-4.8-claude-code
Opus 4.8
claude-code
SWE-bench Verifiedn/a0.0%2302.5 ms
opus-4.8-codex
Opus 4.8
codex
SWE-bench Verifiedn/a0.0%2113 ms
opus-4.8-opencode
Opus 4.8
opencode
SWE-bench Verifiedn/a0.0%2176 ms
opus-4.8-eve
Opus 4.8
eve
SWE-bench Verifiedn/a0.0%2142.5 ms
opus-4.8-mastra
Opus 4.8
mastra
SWE-bench Verifiedn/a0.0%2110 ms