A Mechanistic Investigation of Supervised Fine Tuning

The cosine similarity between a large language model's hidden activations before and after Supervised Fine-Tuning (SFT) remains very high. This, at first glance, suggests that SFT leaves the model's activation geometry largely undisturbed. However, projecting both sets of activations through a Sparse Autoencoder (SAE) pretrained on the base model reveals that the underlying sparse latents diverge significantly. We introduce a novel investigative pipeline which utilizes these pretrained SAEs as a high-resolution diagnostic tool to mechanistically investigate the drivers of this representational divergence. Through our analytical pipeline, we discover task-specific and layer-specific distributions of the precise semantic features that are systematically altered during supervised fine-tuning. We additionally identify a layer-wise update profile specific to safety alignment. All code, experimental scripts, and analysis files associated with this work are publicly available at: https://github.com/ruhzi/sae-investigation.

Paper

References (15)

06. Neuronpedia api documentation2025 · Neuronpedia
082024. Saes (usually) transfer between base and chat modelsAI Alignment Forum / LessWrong
092025. Gemma scope 2 - technical paperarXiv preprint
102024. Openai tool calling dataset (sft-ready)Hugging Face Dataset
11Code (programming constructs, API artifacts, documentation)
12Structure (markdown, bullet points, JSONscaffolding

Scroll for more · 3 remaining

Similar papers

© 2026 NYSGPT2525 LLC