| # | Model | Runner | No Skills Baseline task pass rate when the agent attempts the task without domain skills. This is the leaderboard's rank key. | With Skills Task pass rate when the agent is provided with AtomisticSkills skills and MCP tools. | Skill Boost Absolute pass rate gain: (With-Skills Pass) - (No-Skills Pass). | Time (With / No) Mean wall-clock execution time for With-Skills vs No-Skills trials. | Out Tok (With / No) Mean output tokens generated by the LLM for With-Skills vs No-Skills trials. This represents total completion tokens (usage.completion_tokens) billed by the model provider API, which includes both visible tool-call arguments and internal reasoning/thinking tokens where applicable. | Token Saving Relative reduction in mean output tokens from using skills: (No-Skills β With-Skills) / No-Skills. Positive means the model spent fewer tokens with skills. | Total Cost Total API compute cost for all evaluated runs of this model. | Action |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 |
GPT-5.6 Sol
k=1 (62 runs)
|
codex | 83.9%(26/31) | 96.8%(30/31) | +12.9% | 13m/14m | 11.0k/11.4k | 2.9% |
$76.93
62 runs
|
Explore → |
| 2 |
Gemini 3.7 Flash
k=1 (62 runs)
|
terminus-2 | 80.6%(25/31) | 100.0%(31/31) | +19.4% | 14m/15m | 17.4k/22.9k | 23.9% |
$26.89
62 runs
|
Explore → |
| 3 |
Claude Opus 5
k=1 (62 runs)
|
terminus-2 | 80.6%(25/31) | 100.0%(31/31) | +19.4% | 17m/19m | 10.1k/23.5k | 57.0% |
$58.11
62 runs
|
Explore → |
| 4 |
GLM-5.3 Flash
k=1 (62 runs)
|
terminus-2 | 77.4%(24/31) | 87.1%(27/31) | +9.7% | 32m/41m | 40.8k/63.2k | 35.4% |
$6.68
62 runs
|
Explore → |
| 5 |
DeepSeek V4 Flash
k=1 (62 runs)
|
terminus-2 | 74.2%(23/31) | 87.1%(27/31) | +12.9% | 52m/65m | 60.5k/78.0k | 22.5% |
$10.92
62 runs
|
Explore → |
| 6 |
GPT-5.6 Luna
k=3 (186 runs)
|
codex | 73.1%(68/93) | 95.7%(89/93) | +22.6% | 11m/11m | 11.2k/13.0k | 13.9% |
$14.30
186 runs
|
Explore → |
| 7 |
Qwen3.8-27B
k=1 (62 runs)
|
codex | 51.6%(16/31) | 77.4%(24/31) | +25.8% | 86m/112m | 110.6k/155.1k | 28.7% |
$0.00
62 runs
|
Explore → |
β’ Output Token Accounting: The Out Tok metric reports total completion tokens (usage.completion_tokens / output_tokens) returned by the provider APIs. Under standard provider schemas (OpenAI Responses/Chat, Google GenAI, DeepSeek, Zhipu), this aggregate captures the complete generative loadβsubsuming both emitted tool-call arguments/scripts and internal chain-of-thought/reasoning tokens where active.
β’ Efficiency & Prefix Caching: Multi-turn autonomous coding agents leverage prompt prefix caching (usage.prompt_tokens_details.cached_tokens). Because domain skill documentation (SKILL.md playbooks) is read in early episodes, it stays resident in cache across subsequent turns, enabling >85%β97% cache hit rates at discounted input pricing.
β’ Cost Calculation: Total evaluated cost reflects exact provider API billing rates across all 540 evaluated multi-turn runs (180 for Luna, 60 for each of the other 6 models).