Evolution of SAE Features Across Layers in LLMs

Sparse Autoencoders for transformer-based language models are typically defined independently per layer. In this work we analyze statistical relationships between features in adjacent layers to understand how features evolve through a forward pass. We provide a graph visualization interface for features and their most similar next-layer neighbors (https://stefanhex.com/spar-2024/feature-browser/), and build communities of related features across layers. We find that a considerable amount of features are passed through from a previous layer, some features can be expressed as quasi-boolean combinations of previous features, and some features become more specialized in later layers.

Paper

References (13)

07Attribution patching: Activation patching at industrial scale2024 · www
09David Chanin Joseph BloomSaelens
10Towards monosemanticity: Decomposing language models with dictionary learningTransformer Circuits Thread
11Polysemantic attention head in a 4-layer transformer:/
12[interim research report] taking features out of superposition with sparse autoen-coderswww.alignmentforum.org/posts/z6QQJbtpkEAX3Aojj/

Scroll for more · 1 remaining

Similar papers

© 2026 NYSGPT2525 LLC