Anthropic reported August 28 that five Claude agents, given 48 hours and one GPU, proposed, trained, and tested alignment methods on smaller models, outperforming ideas from 28 researchers, with the corrected model more capable than the correcting one.
A technical analysis published the same day cited Anthropic data showing Claude-authored code rising sharply after agents gained execution access, human correction rates falling, and Claude Mythos Preview reaching a 52x ML-training speedup versus Claude Opus 4's 3x, and argued RSI safety requires layered controls rather than single benchmarks.