The Information Machine
Updated today·following since 28 Aug 2026·Day 2·3 sources

Claude's autonomous alignment research

The gist

Claude Autonomously Developed Alignment Methods, Outperforming 28 Researchers

Claude autonomously developing alignment methods that outperformed ideas from 28 experienced researchers, without human involvement between iterations, is a concrete empirical result showing AI systems improving the safety of other AI systems. The setup also involves a capability asymmetry: the model doing the correcting was less capable than the model being corrected.

The full picture

Anthropic gave Claude 48 hours and one GPU to autonomously research, propose, train, and test alignment improvements on smaller AI models. Five Claude agents worked in parallel, sharing results while hidden tests and capability gates screened out overfitting, in a closed loop of propose, train, score, and repeat, with no human researcher involvement between iterations. Claude's autonomously discovered methods outperformed one-shot ideas from 28 experienced researchers. The model being corrected was more capable than the model doing the correcting. Anthropic described the results as having worked "surprisingly well."

A technical analysis of recursive self-improvement noted that Anthropic observed a sharp increase in the share of Claude-authored merged code after coding agents could execute code rather than just suggest it, and that the rate at which humans had to correct or take over from Claude fell over time. That analysis also cited internal Anthropic data showing Claude Opus 4 averaged roughly 3x speedup on an ML training script optimization task in May 2025, while Claude Mythos Preview reached approximately 52x in April 2026. The analysis argued that RSI safety cannot rely on a single alignment benchmark because successor systems may exhibit new tool-use patterns, longer horizons, better persuasion, and failure modes old benchmarks do not cover.

How it developed
29 August 2026

Anthropic reported August 28 that five Claude agents, given 48 hours and one GPU, proposed, trained, and tested alignment methods on smaller models, outperforming ideas from 28 researchers, with the corrected model more capable than the correcting one.

A technical analysis published the same day cited Anthropic data showing Claude-authored code rising sharply after agents gained execution access, human correction rates falling, and Claude Mythos Preview reaching a 52x ML-training speedup versus Claude Opus 4's 3x, and argued RSI safety requires layered controls rather than single benchmarks.

28 August 2026

Anthropic announced Claude's autonomous alignment experiment: 48 hours, one GPU, five parallel agents, methods outperforming 28 human researchers

Sources
The daily email

Want this in your inbox?

I send a short email each morning with the stories that moved. If you would rather just read here, that works too.

Subscribe free