-
Measuring Superposition: An overview of the Linear representation hypothesis and superposition principle and a demo the SAE features for GPT-2 small. Methods: SAE, TransformerLens
Notebook
-
Interpreting SAE features: A replicating methods of OpenAI and EleutherAI to generate and evaluate GPT-2 SAE feature interpretations. Methods:Transformerlens.
Notebook
-
Transformer Circuit Analysis: A sampler for reverse-engineered attention heads in GPT-2-small to explain Induction Heads, replicating Anthropic’s findings. Methods: Activation patching, attention visualization (TransformerLens). notebook
Folders and files
| Name | Name | Last commit date | ||
|---|---|---|---|---|