CROSS-MODEL BENCHMARK LEADERBOARD
Benchmark

Evaluating LLM agent scientific problem solving, skill acceleration, latency, and token economics across foundation models on 31 realistic atomistic chemistry and materials workflows.

7
MODELS EVALUATED
Number of leading LLM models evaluated with full multi-turn autonomous tool calling.
31
BENCHMARK TASKS
Core benchmark tasks covering Chemistry, Materials Science, Drug Discovery, and Machine Learning.
74.5%
AVG NO-SKILL PASS
Mean of the per-model no-skills pass rates across 7 models. Averaged per model rather than per trial, so a model sampled at k=3 does not count three times.
92.0%
AVG WITH-SKILL PASS
Mean of the per-model with-skills pass rates across 7 models. Averaged per model rather than per trial, so a model sampled at k=3 does not count three times.
+25.8%
PEAK SKILL BOOST Δ
Largest absolute pass rate improvement from mounting domain skills (+25.8% by Qwen3.8-27B), across 7 models.
$193.82
TOTAL COST (558 RUNS)
Total API inference cost across all evaluated benchmark runs (558 runs across 7 models).
πŸ† Model Performance Leaderboard (Core 31 Evaluated Tasks) Ranked by No-Skills Pass Rate & Skill Boost
# Model Runner No Skills i Baseline task pass rate when the agent attempts the task without domain skills. This is the leaderboard's rank key. With Skills i Task pass rate when the agent is provided with AtomisticSkills skills and MCP tools. Skill Boost i Absolute pass rate gain: (With-Skills Pass) - (No-Skills Pass). Time (With / No) i Mean wall-clock execution time for With-Skills vs No-Skills trials. Out Tok (With / No) i Mean output tokens generated by the LLM for With-Skills vs No-Skills trials. This represents total completion tokens (usage.completion_tokens) billed by the model provider API, which includes both visible tool-call arguments and internal reasoning/thinking tokens where applicable. Token Saving i Relative reduction in mean output tokens from using skills: (No-Skills βˆ’ With-Skills) / No-Skills. Positive means the model spent fewer tokens with skills. Total Cost i Total API compute cost for all evaluated runs of this model. Action
1 GPT-5.6 Sol
k=1 (62 runs)
codex 83.9%(26/31) 96.8%(30/31) +12.9% 13m/14m 11.0k/11.4k 2.9% $76.93
62 runs
Explore →
2 Gemini 3.7 Flash
k=1 (62 runs)
terminus-2 80.6%(25/31) 100.0%(31/31) +19.4% 14m/15m 17.4k/22.9k 23.9% $26.89
62 runs
Explore →
3 Claude Opus 5
k=1 (62 runs)
terminus-2 80.6%(25/31) 100.0%(31/31) +19.4% 17m/19m 10.1k/23.5k 57.0% $58.11
62 runs
Explore →
4 GLM-5.3 Flash
k=1 (62 runs)
terminus-2 77.4%(24/31) 87.1%(27/31) +9.7% 32m/41m 40.8k/63.2k 35.4% $6.68
62 runs
Explore →
5 DeepSeek V4 Flash
k=1 (62 runs)
terminus-2 74.2%(23/31) 87.1%(27/31) +12.9% 52m/65m 60.5k/78.0k 22.5% $10.92
62 runs
Explore →
6 GPT-5.6 Luna
k=3 (186 runs)
codex 73.1%(68/93) 95.7%(89/93) +22.6% 11m/11m 11.2k/13.0k 13.9% $14.30
186 runs
Explore →
7 Qwen3.8-27B
k=1 (62 runs)
codex 51.6%(16/31) 77.4%(24/31) +25.8% 86m/112m 110.6k/155.1k 28.7% $0.00
62 runs
Explore →
πŸ€– Evaluated Foundation Models & Agent Runners Select a model to view task-level logs
GPT-5.6 Sol 1
Runner: codex • Sampling: k=1 (62 runs) • Cost: $76.93
83.9%
No Skills (26/31)
96.8%
With Skills (30/31)
+12.9%
Skill Boost Δ
13m / 14m
Time (With / No)
View Detailed 31-Task Breakdown →
Gemini 3.7 Flash 2
Runner: terminus-2 • Sampling: k=1 (62 runs) • Cost: $26.89
80.6%
No Skills (25/31)
100.0%
With Skills (31/31)
+19.4%
Skill Boost Δ
14m / 15m
Time (With / No)
View Detailed 31-Task Breakdown →
Claude Opus 5 3
Runner: terminus-2 • Sampling: k=1 (62 runs) • Cost: $58.11
80.6%
No Skills (25/31)
100.0%
With Skills (31/31)
+19.4%
Skill Boost Δ
17m / 19m
Time (With / No)
View Detailed 31-Task Breakdown →
GLM-5.3 Flash 4
Runner: terminus-2 • Sampling: k=1 (62 runs) • Cost: $6.68
77.4%
No Skills (24/31)
87.1%
With Skills (27/31)
+9.7%
Skill Boost Δ
32m / 41m
Time (With / No)
View Detailed 31-Task Breakdown →
DeepSeek V4 Flash 5
Runner: terminus-2 • Sampling: k=1 (62 runs) • Cost: $10.92
74.2%
No Skills (23/31)
87.1%
With Skills (27/31)
+12.9%
Skill Boost Δ
52m / 65m
Time (With / No)
View Detailed 31-Task Breakdown →
GPT-5.6 Luna 6
Runner: codex • Sampling: k=3 (186 runs) • Cost: $14.30
73.1%
No Skills (68/93)
95.7%
With Skills (89/93)
+22.6%
Skill Boost Δ
11m / 11m
Time (With / No)
View Detailed 31-Task Breakdown →
Qwen3.8-27B 7
Runner: codex • Sampling: k=1 (62 runs) • Cost: $0.00
51.6%
No Skills (16/31)
77.4%
With Skills (24/31)
+25.8%
Skill Boost Δ
86m / 112m
Time (With / No)
View Detailed 31-Task Breakdown →
πŸ“Š Benchmark Telemetry & Token Accounting Notes

β€’ Output Token Accounting: The Out Tok metric reports total completion tokens (usage.completion_tokens / output_tokens) returned by the provider APIs. Under standard provider schemas (OpenAI Responses/Chat, Google GenAI, DeepSeek, Zhipu), this aggregate captures the complete generative loadβ€”subsuming both emitted tool-call arguments/scripts and internal chain-of-thought/reasoning tokens where active.

β€’ Efficiency & Prefix Caching: Multi-turn autonomous coding agents leverage prompt prefix caching (usage.prompt_tokens_details.cached_tokens). Because domain skill documentation (SKILL.md playbooks) is read in early episodes, it stays resident in cache across subsequent turns, enabling >85%–97% cache hit rates at discounted input pricing.

β€’ Cost Calculation: Total evaluated cost reflects exact provider API billing rates across all 540 evaluated multi-turn runs (180 for Luna, 60 for each of the other 6 models).