2024.02.01: Open Source Sparse Autoencoders for all Residual Stream… — ml4se — TG.ME

...
* 2024.02.01: Open Source Sparse Autoencoders for all Residual Stream Layers of GPT2-Small
* 2024.02.06: Challenges in Mechanistically Interpreting Model Representations
* 2024.02.22: Do sparse autoencoders find "true features"?
* 2024.03.14: Sparse autoencoders find composed features in small toy models
* 2024.03.15: Improving SAE's by Sqrt()-ing L1 & Removing Lowest Activating Features
* 2024.03.29: SAE reconstruction errors are (empirically) pathological
* 2024.04.22: Mechanistic Interpretability for AI Safety A Review
* 2024.05.21: Mapping the Mind of a Large Language Mode
* 2024.06.13: The engineering challenges of scaling interpretability
* 2024.07.02: A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models
* 2024.07.29: Detecting and Understanding Vulnerabilities in Language Models via Mechanistic Interpretability
* 2024.10.10: Bilinear MLPs enable weight-based mechanistic interpretability
* 2024.10.11: Explaining AI through mechanistic interpretability
* 2024.10.15: Mechanistic Permutability: Match Features Across Layers
* 2024.10.17: Using Dictionary Learning Features as Classifiers
* 2024.10.24: Probing Ranking LLMs: Mechanistic Interpretability in Information Retrieval
* 2024.10.25: Evaluating feature steering: A case study in mitigating social biases
* 2024.11.25: Adaptive Circuit Behavior and Generalization in Mechanistic Interpretability
👍3
January 9, 2025 941 20