SM
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet
Work
by Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Anthropic Interpretability Team · 2024
This research paper details how sparse autoencoders were scaled to extract millions of interpretable, concept-level features from the production-grade language model Claude 3 Sonnet. By identifying and manipulating these features—such as the 'Golden Gate Bridge' feature—the researchers demonstrated the ability to causally steer the model's behavior and responses.

