The Weight of Silence: A Causal Case for Weights Over the Scratchpad in Latent Chess Reasoning
Latent, or silent, reasoning lets language models carry out intermediate computation in continuous vector space instead of words, and is widely assumed to function as an internal scratchpad that the model actively consults during inference. Whether that assumption survives reinforcement learning has not been tested directly: existing causal analyses of latent reasoning are confined to math and logic tasks, and compare a model's reliance on its thoughts within a single trained checkpoint, never before and after an RL stage. We train a chess-playing model through a staged latent-reasoning curriculum followed by reinforcement learning, and find that legality climbs monotonically to 61% (from a 48% pre-RL baseline) while checkmate confabulation is eliminated entirely. To understand where this gain comes from, we run a six-condition causal intervention suite on the same model before and after reinforcement learning: substituting or adding matched random noise to the latent thought vectors leaves performance unchanged, ablating them (with or without length-matching) causes only mild degradation, and only exact-zero vectors cause collapse. This robustness gap is itself part of the finding: under exact-zero corruption, legality collapses to 1\% pre-RL versus 9\% post-RL, a significant gap that survives correction for testing across the full six-condition battery. Testing the post-RL checkpoint's own side of the battery again at ten times the sample size confirms this pattern holds under substantially more statistical power: removing the thoughts outright and the length-matched removal control both cross into significance as well, in the same severity order the original sample already suggested, while content-preserving substitution and noise remain statistically indistinguishable from undisturbed baseline throughout. Reinforcement learning appears to add robustness to disruption, not reliance on thought content. These results push back against the field's default assumption that latent thoughts function as an actively consulted inference-time scratchpad, and instead indicate that, in this setting, the principal effect of latent reasoning is to shape the model's parameters during training, with much of the downstream improvement encoded in the weights themselves. Alongside this mechanistic result, we demonstrate a working reinforcement-learning gain in chess, a domain outside the math and logic settings where multiple independent groups report the same latent-reasoning-plus-RL recipe failing to improve accuracy over SFT.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex