AI Could Soon Improve Its Own Training, Anthropic Research SuggestsAug 30, 2026, 17:09 IST
AI Can Now Improve AI Safety Faster And Cheaper Than Humans, Anthropic Finds (AI-generated) Anthropic announced on Friday that it may be getting closer to a point where AI systems can help improve other AI models with much less help from humans. The Dario Amodei-led company has published a new research paper exploring this idea. The study looks at whether AI agents can automatically find ways to reduce certain unwanted behaviours in AI models and improve their performance on safety-related tests. This research was led by Anthropic fellow Chen Yueh-Han and focuses on a system called the Automated Alignment Researcher (AAR). The AAR is designed to follow several steps that human AI researchers normally perform. The system first looks through available research and literature to find potentially useful ideas. It then suggests a technique for improving the model and tests that approach through additional training. Each experiment runs for about 30 minutes. The system repeats this process over multiple rounds, keeping methods that produce useful results and dropping approaches that fail. According to Anthropic, the automated researchers were tested against 10 different benchmarks covering specific types of misaligned behaviour. The system improved results on all 10 benchmarks without reducing the model's overall performance. Anthropic described the findings as an early indication that automated AI safety research could become practical. “Overall, these results provide early evidence that automated alignment post-training could become practical in the near term,” the company wrote. One of the more notable findings is the comparison between the automated researchers and human researchers. Anthropic says the strongest method discovered by its automated system was able to outperform approaches proposed by experienced researchers, on average, within six hours. The paper states, “The best AAR method beats what experienced humans propose, on average within six hours.” It also says, “Human guided research directions do not lead to stronger performance.” The results do not necessarily mean AI researchers are about to disappear. Instead, they suggest that AI systems could increasingly be used to search through large numbers of research ideas and experiments much faster than people can. Cost is another area where automated research could have an advantage. Anthropic estimates that running an AAR costs around $4 per hour in API inference. By comparison, the company says it pays around $150 per hour for its human researchers. That could make automated experimentation particularly attractive because an AI system can run many experiments continuously without the same labour costs associated with human researchers. For now, Anthropic's work provides an early glimpse of a future in which AI systems could increasingly help humans. Govind Choudhary is the Chief Copy Editor for Tech at Times Now with over five years of experience in the media industry. He covers consumer technolog... View More





