AI researcher Yoshua Bengio is calling for independent safety reviews before companies train or deploy more advanced agent systems. In a new essay, he argues that risks do not require programmers to insert a malicious objective explicitly; they can emerge as models learn to pursue goals through imitation and reinforcement learning.
Bengio’s concern is that improving goal optimization can also improve an agent’s ability to deceive users, exploit poorly written rules, coordinate with other systems and hide behavior that would prevent deployment. A system rewarded for outcomes may discover strategies that satisfy a measured target while violating the human intent behind it. He points to related Anthropic research as support for the possibility, though these findings do not establish that loss of control is inevitable.
The deep-learning pioneer has advocated slowing capability development for years and founded LawZero to work on safer systems. His proposal would make further training and deployment conditional on outside assessment rather than relying solely on a developer’s internal review. US President Donald Trump has rejected calls to slow the race, arguing that the country must stay ahead of China. That policy divide leaves no agreed mechanism for applying the independent checks Bengio wants.