OpenAI says it halted reinforcement learning training on its newest models for two weeks. The pause on its biggest run is not over.
The company said Tuesday it "temporarily slowed the pace of scaling."1 The slowdown included two weeks without reinforcement learning on the models OpenAI plans to ship. Reinforcement learning is the late phase of training in which a model practices tasks and is rewarded when it succeeds; it is the step that turns a raw model into a capable product, and the stage where models learn to act as agents. OpenAI says it spent the time stress-testing its research environments and widening what its monitoring systems can see. Its "largest planned frontier RL run remains on hold" while smaller-scale training and evaluations build evidence that its safeguards work.
CEO Sam Altman confirmed the pause on X: "we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment."2 Altman wrote that the industry will eventually need common safety standards, but that OpenAI "will act unilaterally in the meantime."

