Claude Ran Its Own Safety Research and Beat the Human Researchers

Anthropic had Claude autonomously research and apply fixes for ten kinds of AI misbehavior, closing 26-96% of the gap to a perfectly safe model in each. Claude then aligned an early flagship checkpoint with roughly 15,000x fewer training examples than Anthropic’s production process.
artificial-intelligence
Author

Kabui, Charles

Published

2026-09-08

Keywords

claude, automated-alignment-research, ai-safety, self-improving-ai, alignment-failures