Anthropic Researcher Offers a Peek at Self-Improving AI

A new paper shows automated systems reliably fixing alignment failures — and beating human researchers on cost

Published: 2026-08-29 Category: Quick Take Sources: TechCrunch

The automated alignment researcher

Training AI models with other AI models has become a popular goal across the frontier labs, and a researcher in Anthropic's fellows program has now given us an early look at what it looks like in practice. On Friday, Anthropic published a paper titled "Automated Researchers Can Reliably Mitigate Alignment Failures," detailing how AI systems could reliably improve a model's performance on a set of alignment benchmarks. Given 10 benchmarks for specific misaligned behaviors, the automated systems improved performance on every single one without degrading overall performance.

Led by Anthropic fellow Chen Yueh-Han, the system replicates much of the traditional research approach. Each automated system searches the available literature, proposes a method, and trains the model using that method for 30 minutes, gradually increasing the benchmark over several iterations. Effective methods are preserved while ineffective ones are discarded, letting the system operate quickly and at scale.

A step toward recursive self-improvement

The paper is a step toward recursive self-improvement, which many see as the next significant milestone in AI progress. If models can improve their own alignment training, the logic goes, they could plausibly improve training practices more broadly — at which point human AI researchers might soon become obsolete. The paper isn't shy about that implication, explicitly comparing the Automated Alignment Researcher (AAR) to its human equivalent.

"The best AAR method beats what experienced humans propose, on average within six hours," the paper reads. "Human guided research directions do not lead to stronger performance." There's even a cost comparison: "An AAR costs roughly $4 per hour in API inference against the $150 per hour we pay our human researchers."

The caveats

The paper is careful to flag limitations. The automated system only works insofar as the benchmarks reflect the actual alignment goals, and there's significant work to be done in establishing and maintaining those benchmarks — not to mention maintaining and expanding the literature the automated researchers draw from. In other words, the machine is only as good as the human scaffolding around it, at least for now.

Still, the direction is unmistakable. The economics alone — $4 an hour versus $150 — make this an irresistible path for labs under pressure to scale safety work. The question is no longer whether models will help train models, but how quickly the human loop shrinks to just defining the benchmarks.

Source: TechCrunch, "An Anthropic researcher just gave us a peek at self-improving AI" (Russell Brandom, August 28, 2026).