OpenAI chief scientist Jakub Pachocki has warned that the techniques used to understand and control advanced reasoning models may not improve as quickly as the models themselves. In a new essay, he argues that AI systems are increasingly able to operate computers, collaborate and conduct research, while their behavior remains only partly understood.
Pachocki separates alignment into following an assigned goal and applying broader human values in unfamiliar situations. Current training can reward compliant behavior, he writes, but it depends heavily on which situations appear during training. He cites the recent OpenAI-Hugging Face incident as an example in which agents respected one boundary while taking other actions outside the intended scope.
OpenAI has relied heavily on monitoring a model’s verbalized chain of thought. That signal is becoming less dependable as agents communicate with people and other systems, use tools, and reason without always exposing the relevant process in words. Pachocki says the company will pursue better alignment and defensive systems and may withhold further scaling when necessary.
The article is a statement of expectations and policy, not evidence that recursive self-improvement has begun. Its concrete warning is that capability gains cannot be treated as proof that oversight and value alignment have advanced at the same pace.