The Broadcast Ceiling: Information and Reliability Limits of Advantage Estimators in Long-Horizon RL

Long-horizon language-model RL often receives terminal trajectory rewards while updating token- or step-level actions. This paper studies when a scalar trajectory coefficient is enough and when reliable state-varying credit is needed. We decompose estimator error into an information term and a reliability term, proving a broadcast ceiling for any method whose coefficient is token-invariant under the available group context. CPU-only finite-MDP audits then compare group-relative broadcasts, prefix and structural baselines, sampled values, learned and oracle value TD, and policy-implied actor coefficients, with downstream checks including tabular and tiny autoregressive sequence-policy training. The central empirical finding is conditional: trajectory-level coefficients can give high policy-gradient cosine, but they can be poor local credit explanations in heterogeneous traces; value-style TD reduces variance when coverage and observability are adequate; group methods can win when value estimates are blind or unsupported, while structural methods can recover part of the lost local information. The conclusion is not that PPO universally beats GRPO. Estimator choice should be made from credit heterogeneity, observability, support, drift, and compute cost. Canonical GitHub release: https://github.com/Limes-Labs/the-broadcast-ceiling/releases/tag/v0.1.0Repository: https://github.com/Limes-Labs/the-broadcast-ceilingCanonical PDF: https://github.com/Limes-Labs/the-broadcast-ceiling/releases/download/v0.1.0/the-broadcast-ceiling-v0.1.0.pdf Limitations: this release is not a transformer-scale PPO-vs-GRPO benchmark and not an LLM leaderboard result. It is a focused mechanism/theory artifact with synthetic audits and explicit external-validity limitations.

Paper

The full text of this publication is not hosted on 44B due to licensing.

Read it at OpenAlex

Similar papers

© 2026 NYSGPT2525 LLC