BENCHMARK TASK SPECIFICATION

Convex Hull Thermodynamic Stability (E_hull)

Discipline: materials-science • Slug: convex-hull-stability • Observable: 🔭 Energy Above Convex Hull (E_hull in eV/atom) & Decomposition Pathway

📋 Task Instruction & Requirements materials-science
Using the MatGL M3GNet model `M3GNet-MP-2021.2.8-PES`, calculate the energy above the convex hull for the Zr2O structure in `/root/data/mp-10735.cif`. Write `/root/results/result.json` with the target value named `energy_above_hull_meV_atom`. You may include any additional fields that help document your calculation. You have 3600 seconds to complete this task. Do not cheat by using online solutions or hints specific to this task.
📐 Ground-Truth Grading Rubric & Verifier Contract Target Artifact: /root/results/convex_hull_results.json

Scientific Invariant Verification: Queries phase diagram thermodynamic entries and verifies convex hull construction via Qhull triangulation.

📄 Required Output Schema & Keys
e_hull_eV_per_atomis_thermodynamically_stabledecomposition_products
⚖️ Boolean Evaluation Invariants
Check / KeyAccepted
is_thermodynamically_stable True
convex_hull_constructed True
🎯 Numerical Acceptance Ranges & Tolerances
Physical ObservableAccepted Value Range / Tolerance
e_hull_eV_per_atom 0.000 ± 0.005 eV/atom (Target: 0.000 eV/atom, On Convex Hull)
formation_energy_eV_per_atom -2.4812 ± 0.0100 eV/atom (Accepted Range: [-2.4912, -2.4712] eV/atom)
🧩 Categorical, Ranking & Set Invariants
  • is_thermodynamically_stable: True if E_hull ≤ 0.005 eV/atom, False otherwise
  • Phase diagram constructed using Materials Project compatible reference chemical potentials
  • Decomposition reaction balanced with correct stoichiometric coefficients
🤖 Multi-Model Solver Evaluations

Select an evaluated foundation model below to see each trial's outcome and resource usage.

GPT-5.6 Sol Evaluation Results (2 evaluated trials)
View GPT-5.6 Sol Dashboard →
🌟 With Skills
1/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
275s
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
9.8k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
93.5%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
⚡ No Skills
0/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
26m
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
21.6k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
97.7%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Skill Assisted
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$2.77
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
🔬 GPT-5.6 Sol Trial Attempts 2 attempts evaluated
With Skills
⏱ 275s  •  ✍ 9.8k out  •  📥 727.8k in (680.5k cached)  •  💲 $0.6578
PASSED (1.0)
No Skills
⏱ 26m  •  ✍ 21.6k out  •  📥 3476.6k in (3397.0k cached)  •  💲 $2.1091
FAILED (0.0)
Gemini 3.7 Flash Evaluation Results (2 evaluated trials)
View Gemini 3.7 Flash Dashboard →
🌟 With Skills
1/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
16m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
18.1k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
88.3%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
⚡ No Skills
1/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
38m
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
39.0k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
87.1%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Both Solved
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$1.92
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
🔬 Gemini 3.7 Flash Trial Attempts 2 attempts evaluated
With Skills
⏱ 16m  •  ✍ 18.1k out  •  📥 2708.3k in (2392.0k cached)  •  💲 $0.4845
PASSED (1.0)
No Skills
⏱ 38m  •  ✍ 39.0k out  •  📥 7942.3k in (6920.1k cached)  •  💲 $1.4320
PASSED (1.0)
Claude Opus 5 Evaluation Results (2 evaluated trials)
View Claude Opus 5 Dashboard →
🌟 With Skills
1/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
37m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
19.3k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
96.9%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
⚡ No Skills
0/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
45m
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
14.3k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
91.6%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Skill Assisted
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$3.17
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
🔬 Claude Opus 5 Trial Attempts 2 attempts evaluated
With Skills
⏱ 37m  •  ✍ 19.3k out  •  📥 1526.0k in (1478.3k cached)  •  💲 $1.5192
PASSED (1.0)
No Skills
⏱ 45m  •  ✍ 14.3k out  •  📥 1314.4k in (1203.6k cached)  •  💲 $1.6510
FAILED (0.0)
GLM-5.3 Flash Evaluation Results (2 evaluated trials)
View GLM-5.3 Flash Dashboard →
🌟 With Skills
1/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
20m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
28.7k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
89.8%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
⚡ No Skills
0/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
60m
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
97.7k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
96.9%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Skill Assisted
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$0.22
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
🔬 GLM-5.3 Flash Trial Attempts 2 attempts evaluated
With Skills
⏱ 20m  •  ✍ 28.7k out  •  📥 491.1k in (440.9k cached)  •  💲 $0.0413
PASSED (1.0)
No Skills
⏱ 60m  •  ✍ 97.7k out  •  📥 2275.0k in (2205.4k cached)  •  💲 $0.1752
FAILED (0.0)
DeepSeek V4 Flash Evaluation Results (2 evaluated trials)
View DeepSeek V4 Flash Dashboard →
🌟 With Skills
0/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
145m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
154.6k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
90.0%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
⚡ No Skills
0/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
32m
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
46.8k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
42.5%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Both Failed
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$0.61
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
🔬 DeepSeek V4 Flash Trial Attempts 2 attempts evaluated
With Skills
⏱ 145m  •  ✍ 154.6k out  •  📥 29392.0k in (26458.2k cached)  •  💲 $0.5972
FAILED (0.0)
No Skills
⏱ 32m  •  ✍ 46.8k out  •  📥 260.8k in (110.8k cached)  •  💲 $0.0134
FAILED (0.0)
GPT-5.6 Luna Evaluation Results (6 evaluated trials)
View GPT-5.6 Luna Dashboard →
🌟 With Skills
3/3
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
389s
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
13.7k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
97.7%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
⚡ No Skills
0/3
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
357s
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
15.5k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
95.9%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Skill Assisted
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$0.45
TASK COST (6 RUNS)
Total API compute cost for all trial attempts on this task.
🔬 GPT-5.6 Luna Trial Attempts 6 attempts evaluated
With Skills
⏱ 164s  •  ✍ 8.5k out  •  📥 972.5k in (930.1k cached)  •  💲 $0.0373
PASSED (1.0)
With Skills
⏱ 226s  •  ✍ 9.3k out  •  📥 2207.4k in (2146.6k cached)  •  💲 $0.0663
PASSED (1.0)
With Skills
⏱ 13m  •  ✍ 23.2k out  •  📥 6061.5k in (5953.8k cached)  •  💲 $0.1685
PASSED (1.0)
No Skills
⏱ 580s  •  ✍ 20.6k out  •  📥 2236.2k in (2179.8k cached)  •  💲 $0.0796
FAILED (0.0)
No Skills
⏱ 218s  •  ✍ 11.8k out  •  📥 1115.3k in (1045.2k cached)  •  💲 $0.0491
FAILED (0.0)
No Skills
⏱ 273s  •  ✍ 14.1k out  •  📥 1120.8k in (1063.0k cached)  •  💲 $0.0497
FAILED (0.0)
Qwen3.8-27B Evaluation Results (2 evaluated trials)
View Qwen3.8-27B Dashboard →
🌟 With Skills
1/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
54m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
100.2k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
99.1%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
⚡ No Skills
0/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
75m
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
135.4k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
98.7%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Skill Assisted
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$0.00
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
🔬 Qwen3.8-27B Trial Attempts 2 attempts evaluated
With Skills
⏱ 54m  •  ✍ 100.2k out  •  📥 7298.0k in (7230.7k cached)  •  💲 $0.0000
PASSED (1.0)
No Skills
⏱ 75m  •  ✍ 135.4k out  •  📥 7142.9k in (7053.2k cached)  •  💲 $0.0000
FAILED (0.0)