External registration
Register externallyPaper Discussion on Mechanistic Interpretability - Identifying Human-Interpretable Concepts and Algorithms in LLMs
Read some of the AI headlines from the last weeks. If you still think that interpretability and improving our understanding of LLMs is not important, read them again.
In general, mechanistic interpretability aims to discover, understand and verify the algorithms that model weights implement by reverse engineering model computation into human-understandable components. To be specific, we want to look at the Interpretability in the Wild paper in this session (https://openreview.net/pdf?id=NpsVSN6o4ul It identifies a circuit for indirect object identification in the GPT-2 small model, and has been an influential real-world proof of concept in the field.
Before the session, please have a read of the paper (https://openreview.net/pdf?id=NpsVSN6o4ul The more questions you have after reading, the more productive the session will be! You can also have a look at this more hands-on interpretability course chapter on the same paper: https://learn.arena.education/chapter1_transformer_interp/21_ioi/intro