SM

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

Work

by Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Anthropic Interpretability Team · 2024

This research paper details how sparse autoencoders were scaled to extract millions of interpretable, concept-level features from the production-grade language model Claude 3 Sonnet. By identifying and manipulating these features—such as the 'Golden Gate Bridge' feature—the researchers demonstrated the ability to causally steer the model's behavior and responses.

1 episode featuring Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

What is PodLume?

PodLume turns podcasts into searchable knowledge. AI-decoded transcripts, identified guests and topics, smart highlights, and cross-show search across the world’s best conversations — all in your pocket.

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet | PodLume