Leaderboard
Scoring Methodology¶
- Scores are macro-averaged over the 60 tasks and, unless noted otherwise, over three independent runs; the website does not recompute them from task-level artifacts.
- Nineteen results report the mean over three runs. Claude Opus 5 and Qwen3.8-Max are the single-run exceptions and are labeled accordingly.
- The published evaluation was conducted without external tool access.
- The evaluation task count is 60. A scored task-set identifier was not supplied, so the exact task-set version remains explicitly unverified.
- The score snapshot is kept separate from the public task catalog, which is independently generated from a pinned public repository revision.
- Kimi Code + Kimi K3, Claude Code + DeepSeek V4 Flash, and OpenHands + DeepSeek V4 Flash appear in this supplied ranking; Claude Fable 5 does not appear in the latest source.
- B1 provides the complete method and procedure; B2 retains the method but removes procedural detail; B3 requires the system to determine the method; and B4 adds factually correct but non-essential context to B3.
Across all 21 configurations, the mean score is 51.72 / 29.72 / 27.48 / 27.63 for B1–B4. The largest decline occurs from B1 to B2 (−22.01), indicating that method operationalization is a larger current bottleneck than method selection or distractor robustness.
Results should only be compared directly after confirming the same benchmark release, task set, scorer, runner, and runtime policy.
Data Availability¶
The current aggregate snapshot contains 21 Agent×Model configuration rows with overall and B1–B4 means plus the adopted source rounds. Qwen3.8-Max, Gemini 3.8 Flash, and GLM-5.3-Flash are listed under Claude Code with Max effort and linked to official release announcements.
Once task-level reporting artifacts are available, they can be added as drill-down details from each result without creating another top-level leaderboard page.