AgentCompass: 21 Agent Benchmarks Through One Evaluation Layer

AgentCompass separates benchmarks, agent harnesses, models, and environments so researchers can compare runs without rebuilding each test. Its paper found suspected reward hacking in up to 39.12% of correct coding samples.
artificial-intelligence
software-engineering
Author

Kabui, Charles

Published

2026-07-26

Keywords

agentcompass, agent-evaluation, llm-benchmarks, reward-hacking, reproducible-evaluation