Navigating the Concept Space of Language Models

Sparse autoencoders (SAEs) trained on large language model activations output thousands of features that enable mapping to human-interpretable concepts. The current practice for analyzing these features primarily relies on inspecting top-activating examples, manually browsing individual features, or performing semantic search on interested concepts, which makes exploratory discovery of concepts difficult at scale. In this paper, we present Concept Explorer, a scalable interactive system for post-hoc exploration of SAE features that organizes concept explanations using hierarchical neighborhood embeddings. Our approach constructs a multi-resolution manifold over SAE feature embeddings and enables progressive navigation from coarse concept clusters to fine-grained neighborhoods, supporting discovery, comparison, and relationship analysis among concepts. We demonstrate the utility of Concept Explorer on SAE features extracted from SmolLM2, where it reveals coherent high-level structure, meaningful subclusters, and distinctive rare concepts that are hard to identify with existing workflows.

Paper

References (22)

0919336 tokens wrapped in double angle brackets — especially the placeholder ” << ” . ” >> ” used to obfuscate periods in URLs/domain nameshelp.syncfusion <
10EleutherAIEleutherai/sae-smollm2-135m-64x
1116758 tokens enclosed in double angle brackets ( <<>> ) — i.e., highlighted/- target words — responding to the formatting cue (across parts of speech and semantics) rather than a specific lexical meaning
12Towards monosemanticity: Decomposing language models with dictionary learningTransformer Circuits Thread

Scroll for more · 10 remaining

Similar papers

© 2026 NYSGPT2525 LLC