The repository implements a system for interpreting language mechanisms using sparse autoencoders, named SAELing. The system aims to reveal and control the internal linguistic knowledge of large language models. We use SAELing to extract a large number of causal features from large language models. For details, see Sparse Auto-Encoder Interprets Linguistic Features in Large Language Models.
The Camera Ready Version Code Repo is Here. Any updates will be maintained on the new repository.