Emergent Depthwise Activation Structure in Decoder-Only Transformer Language Models: The Hau Curve and Its Early-Training Convergence Pattern
Decoder-only transformer language models (LLMs) exhibit a highly regular internal organization in how computation is allocated across depth, across the model families, checkpoints, and scales examined. In this study, we identify a robust, tri-phasic depthwise activation geometry summarized by three operational landmarks: the Early Layer Extremum (πΈπ₯), Mid-layer Plateau (ππ), and Late Layer Surge (πΏπ ). By analyzing layer-wise β1 activations across a 60-model core cohort (augmented with 17 additional models for boundary mapping, total πππβππ π‘ = 77), we show that these landmarks emerge early in pretraining (within roughly 7β15% of steps in the examined runs) and then remain stable. Crucially, the available evidence is consistent with an emergence zone centered around ππππ¦ππ π β 8β12 layers rather than a sharp universal threshold. Below this approximate range, the proximity of the πΈπ₯ and the πΏπ leads to phase congestion, where limited depth resolution causes early- and late-phase activation regimes to overlap and suppress clear expression of the ππ. Above this range, the πΈπ₯ and the πΏπ decouple sufficiently to provide the computational real estate for a stabilized ππ. Together, these results suggest that increasing decoder depth is not merely a quantitative increase in parameters, but also an expansion of the geometric room available for a recurring activation structure. These structural regularities offer a new lens for model analysis, bridging the gap between low-level mechanistic interpretability and high-level behavioral scaling laws. Notably, quantities such as layerwise activation magnitudeβoften treated as secondary or unstableβare shown here to track a recurring architecture-level geometric structure that emerges early and persists across sufficiently deep decoder-only models.
Paper
The full text of this publication is not hosted on 44B due to licensing.
Read it at OpenAlex