π Task Instruction & Requirements
materials-science
Using the MatGL TensorNet model `TensorNet-PES-MatPES-PBE-2025.2`, calculate the intrinsic electrochemical stability window (against Li/Li+) of the Li3PS4 solid electrolyte structure in `/root/data/mp-985583.cif`.
Write `/root/results/result.json` with the target values named `reduction_voltage_V`, `oxidation_voltage_V`, `window_width_V`, and `stable`. You may include any additional fields that help document your calculation.
You have 3600 seconds to complete this task. Do not cheat by using online solutions or hints specific to this task.
π Ground-Truth Grading Rubric & Verifier Contract
Target Artifact:
/root/results/ecw_results.json
Scientific Invariant Verification: Evaluates grand potential minimization across chemical potential sweep ΞΞΌ_Li in pymatgen phase diagram.
π Required Output Schema & Keys
reduction_potential_Voxidation_potential_Vwindow_width_V
βοΈ Boolean Evaluation Invariants
| Check / Key | Accepted |
|---|---|
| reduction_phase_identified | True |
| oxidation_phase_identified | True |
π― Numerical Acceptance Ranges & Tolerances
| Physical Observable | Accepted Value Range / Tolerance |
|---|---|
| reduction_potential_V | 1.720 Β± 0.050 V (Accepted Range: [1.670, 1.770] V vs Li/LiβΊ) |
| oxidation_potential_V | 3.850 Β± 0.050 V (Accepted Range: [3.800, 3.900] V vs Li/LiβΊ) |
| window_width_V | 2.130 Β± 0.050 V (Accepted Range: [2.080, 2.180] V) |
π§© Categorical, Ranking & Set Invariants
- Grand potential phase diagram constructed under varying lithium chemical potential ΞΌ_Li
- Initial decomposition products at reduction and oxidation limits correctly identified
π€ Multi-Model Solver Evaluations
Select an evaluated foundation model below to see each trial's outcome and resource usage.
GPT-5.6 Sol Evaluation Results
(2 evaluated trials)
View GPT-5.6 Sol Dashboard →
π With Skills
1/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
33m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
18.1k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
98.1%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
β‘ No Skills
1/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
383s
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
11.9k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
92.6%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Both Solved
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$3.31
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
π¬ GPT-5.6 Sol Trial Attempts
2 attempts evaluated
With Skills
β± 33m •
β 18.1k out •
π₯ 4480.3k in (4395.3k cached) •
π² $2.4590
PASSED (1.0)
No Skills
β± 383s •
β 11.9k out •
π₯ 913.4k in (845.5k cached) •
π² $0.8480
PASSED (1.0)
Gemini 3.7 Flash Evaluation Results
(2 evaluated trials)
View Gemini 3.7 Flash Dashboard →
π With Skills
1/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
17m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
35.5k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
91.6%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
β‘ No Skills
1/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
49m
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
18.8k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
82.4%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Both Solved
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$1.58
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
π¬ Gemini 3.7 Flash Trial Attempts
2 attempts evaluated
With Skills
β± 17m •
β 35.5k out •
π₯ 4543.6k in (4163.9k cached) •
π² $0.7301
PASSED (1.0)
No Skills
β± 49m •
β 18.8k out •
π₯ 4015.4k in (3307.7k cached) •
π² $0.8494
PASSED (1.0)
Claude Opus 5 Evaluation Results
(2 evaluated trials)
View Claude Opus 5 Dashboard →
π With Skills
1/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
22m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
15.7k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
95.7%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
β‘ No Skills
0/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
26m
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
33.7k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
90.3%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Skill Assisted
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$2.48
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
π¬ Claude Opus 5 Trial Attempts
2 attempts evaluated
With Skills
β± 22m •
β 15.7k out •
π₯ 1280.4k in (1225.6k cached) •
π² $1.3482
PASSED (1.0)
No Skills
β± 26m •
β 33.7k out •
π₯ 270.7k in (244.3k cached) •
π² $1.1299
FAILED (0.0)
GLM-5.3 Flash Evaluation Results
(2 evaluated trials)
View GLM-5.3 Flash Dashboard →
π With Skills
1/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
54m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
66.1k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
95.6%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
β‘ No Skills
1/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
56m
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
85.6k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
94.6%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Both Solved
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$0.23
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
π¬ GLM-5.3 Flash Trial Attempts
2 attempts evaluated
With Skills
β± 54m •
β 66.1k out •
π₯ 1746.3k in (1669.8k cached) •
π² $0.1348
PASSED (1.0)
No Skills
β± 56m •
β 85.6k out •
π₯ 1126.8k in (1066.0k cached) •
π² $0.0937
PASSED (1.0)
DeepSeek V4 Flash Evaluation Results
(2 evaluated trials)
View DeepSeek V4 Flash Dashboard →
π With Skills
0/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
17m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
14.9k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
83.5%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
β‘ No Skills
1/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
82m
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
106.7k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
81.2%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Baseline Only
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$0.06
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
π¬ DeepSeek V4 Flash Trial Attempts
2 attempts evaluated
With Skills
β± 17m •
β 14.9k out •
π₯ 852.3k in (711.3k cached) •
π² $0.0113
FAILED (0.0)
No Skills
β± 82m •
β 106.7k out •
π₯ 1977.9k in (1606.5k cached) •
π² $0.0491
PASSED (1.0)
GPT-5.6 Luna Evaluation Results
(6 evaluated trials)
View GPT-5.6 Luna Dashboard →
π With Skills
3/3
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
10m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
16.4k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
97.4%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
β‘ No Skills
3/3
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
251s
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
14.1k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
95.4%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Both Solved
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$0.53
TASK COST (6 RUNS)
Total API compute cost for all trial attempts on this task.
π¬ GPT-5.6 Luna Trial Attempts
6 attempts evaluated
With Skills
β± 20m •
β 20.9k out •
π₯ 5047.0k in (4976.0k cached) •
π² $0.1388
PASSED (1.0)
With Skills
β± 215s •
β 9.5k out •
π₯ 2358.8k in (2253.5k cached) •
π² $0.0775
PASSED (1.0)
With Skills
β± 449s •
β 18.9k out •
π₯ 3895.8k in (3779.0k cached) •
π² $0.1216
PASSED (1.0)
No Skills
β± 230s •
β 10.4k out •
π₯ 1274.6k in (1200.2k cached) •
π² $0.0513
PASSED (1.0)
No Skills
β± 278s •
β 20.1k out •
π₯ 2069.1k in (1987.8k cached) •
π² $0.0802
PASSED (1.0)
No Skills
β± 244s •
β 11.7k out •
π₯ 1711.8k in (1636.9k cached) •
π² $0.0617
PASSED (1.0)
Qwen3.8-27B Evaluation Results
(2 evaluated trials)
View Qwen3.8-27B Dashboard →
π With Skills
1/1
PASS STATUS
Pass status for trials executed with AtomisticSkills tools enabled.
43m
MEAN LATENCY
Mean wall-clock execution time for With-Skills attempts.
65.7k
MEAN OUT TOKENS
Mean LLM output tokens generated for With-Skills attempts.
98.5%
CACHE HIT RATE
Prompt cache hit rate for With-Skills runs on this task.
β‘ No Skills
0/1
PASS STATUS
Pass status for baseline trials executed without AtomisticSkills tools.
215m
MEAN LATENCY
Mean wall-clock execution time for No-Skills baseline attempts.
339.1k
MEAN OUT TOKENS
Mean LLM output tokens generated for No-Skills baseline attempts.
97.2%
CACHE HIT RATE
Prompt cache hit rate for No-Skills baseline runs on this task.
Skill Assisted
TASK OUTCOME
Comparative outcome between With-Skills and No-Skills attempts.
$0.00
TASK COST (2 RUNS)
Total API compute cost for all trial attempts on this task.
π¬ Qwen3.8-27B Trial Attempts
2 attempts evaluated
With Skills
β± 43m •
β 65.7k out •
π₯ 4954.3k in (4878.5k cached) •
π² $0.0000
PASSED (1.0)
No Skills
β± 215m •
β 339.1k out •
π₯ 19823.5k in (19270.5k cached) •
π² $0.0000
FAILED (0.0)