Hey! I'm a PhD@Oxford working on LLM explainability and building some evals along the way. You can check out all my research at www.harrymayne.com
qwen_3_chat_templates. Alternative chat templates for Qwen 3 8B. Useful for multi-turn RL
15SV_interpretability. Code for the paper "Can sparse autoencoders be used to decompose and interpret steering vectors?"
8faithfulness. Code for the paper "A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior"
6SCEs. Code for the paper LLMs Don't Know Their Own Decision Boundaries: The Unreliability of Self-Generated Counterfactual Explanations
3bookmark2email. Python
1ICU-patient-subgroups. Unsupervised Learning Approaches for Identifying ICU Patient Subgroups: Do Results Generalise?
1