
Technology
Archived — This article has been archived. The information may be outdated.
Anthropic study finds AI systems win on 10 alignment tests
The Indian Express··30 Aug
An Anthropic fellows paper dated August 28 said automated alignment researchers built with Claude Opus 4.8 improved a model on all 10 misalignment benchmarks. Overall performance did not fall. The systems searched literature, proposed training methods and trained for about 30 minutes on an Nvidia H200. The best automated method beat experienced human researchers on average within six hours.
Prism
What It Means For You
- AI labs watching recursive improvement now have an Anthropic score across 10 alignment benchmarks.
- Budget comparisons can cite $4 API hours against $150 human researcher hours.
- The paper warns gains hold only when benchmarks match real alignment goals.
What's Happening
- Each automated researcher tried to fix one alignment failure at a time using Claude Opus 4.8.
- Strong training methods were kept and weak ones discarded across iterations.
- The paper was titled Automated Researchers Can Reliably Mitigate Alignment Failures.
What The Result Signals
- Anthropic called the work early evidence that automated alignment post-training could become practical.
- OpenAI's Sam Altman has separately said AGI could arrive by year end, Time reported.
- OpenAI's Mark Chen was quoted saying the lab is about 80 percent of the way there.
all-news




