ScienceIDE: A Training Environment Where AI Agents Are Graded by Physics

PhAI Labs turned 27 scientific simulation programs into 64 practice environments and 2,812 repair tasks, where an AI agent is rewarded only if the simulation reproduces its reference numbers. Training on those tasks lifted held-out reward from 0.36 to 0.86.
artificial-intelligence
data-science
Author

Kabui, Charles

Published

2026-09-22

Keywords

scientific-code, agent-environments, reinforcement-learning, verifiable-rewards, code-repair