Prior research has shown that humans possess an internal world model that lets them simulate actions and their effects on the state of the world, enabling deliberate planning for complex tasks such as motor control, imagination, reasoning, and decision making. LLMs, by contrast, can only reason autoregressively. To close this gap, the authors bring reinforcement learning and Monte Carlo tree search into LLM reasoning.