Awesome-Interpretability-in-Large-Language-Models. This repository collects all relevant resources about interpretability in LLMs
402HR-VAE. Code for the paper "A Stable Variational Autoencoder for Text Modelling"
26ARC_JSD. A Jensen-Shannon Divergence Driven Mechanistic Study of Context Attribution in Retrieval-Augmented Generation
15TWR-VAE. Improving Variational Autoencoder for Text Modelling withTimestep-Wise Regularisation
5Anchored_Bias_GPT2. This is the official code for anchored bias of GPT2
4