Anthropic has published a new paper outlining how AI systems can improve their own alignment training, a step toward what researchers call recursive self-improvement. The paper, titled “Automated Researchers Can Reliably Mitigate Alignment Failures,” describes a system that improved performance on ten separate alignment benchmarks without degrading overall model quality. The work was led by Chen Yueh-Han, a researcher in Anthropic’s fellows program.

The automated system mimics a typical human research workflow, according to the paper. Each AI agent searches existing literature, proposes a method, and then trains a model using that approach for 30 minutes, iterating across several rounds to boost its benchmark score. Successful methods are kept, while ineffective ones are discarded, allowing the system to operate quickly and at a larger scale than human-led efforts. The paper states that these results offer early evidence that automated alignment post-training could become practical in the near term.

The research addresses the prospect of AI systems improving their own training procedures, a development many in the field see as a key next milestone. The paper directly compares its automated alignment researcher, or AAR, to human counterparts, noting that the best AAR method outperforms what experienced human researchers propose, on average, within six hours. It adds that human-guided research directions did not lead to stronger performance in the tests.

Cost is also a factor in the comparison. The paper states that running the automated researcher costs roughly $4 per hour in API inference, compared with the $150 per hour Anthropic pays its human researchers. This cost difference is presented as part of the case for automated approaches, though the paper also acknowledges limits to the system’s utility.

The authors note that the automated method is only as effective as the benchmarks it is trained against, and that maintaining those benchmarks remains significant work. They also point to the need for ongoing expansion of the research literature the automated system draws from. The paper does not claim that human researchers will soon be obsolete, but it does raise the question of how much of the research process can be automated.

The findings come as several AI labs pursue training models with other models, an area of growing interest in the industry. Anthropic’s paper offers one of the first detailed looks at how that approach could apply to alignment specifically. The paper stops short of saying when such systems might be deployed broadly, instead framing the results as an early demonstration of a method that could become practical in the near term.

More AI news from TechManNews.