Automated research agents improved an AI model on ten tests of unwanted behavior without reducing its overall benchmark performance, according to a paper published by Anthropic. The experiment offers a narrow demonstration of models helping to design their own post-training rather than evidence of unrestricted self-improvement.
Each Automated Alignment Researcher searched available literature, proposed a method and trained the target model for 30 minutes. Successful ideas were retained while weaker ones were discarded over repeated rounds. The paper reports that the best automated method surpassed proposals from experienced human researchers within six hours. It estimates inference cost at about $4 an hour, compared with the study’s $150 hourly rate for human researchers.
The result depends heavily on the tests chosen. An agent can optimize a benchmark that incompletely represents the desired behavior, and people still have to create, maintain and interpret those evaluations and the research literature used by the system. Performance on ten defined failure modes does not show that a model became broadly safer or can discover every hidden weakness. The practical near-term use is more limited: automation may let alignment teams test many post-training ideas cheaply and consistently, while humans remain responsible for deciding whether the measured target reflects the real safety goal.