chess_llm_interpretability. Visualizing the internal board state of a GPT trained on chess PGN strings, and performing interventions on its internal board state and representation of player Elo.
220SAEBench. Python
179activation_oracles. Python
95chess_gpt_eval. A repo to evaluate various LLM's chess playing abilities.
91train_ChessGPT. A repository for training nanogpt-based Chess playing language models.
30dictionary_learning_demo. Jupyter Notebook
26SAE_BoardGameEval. Python
25difr. Python
6llm_bias. Python
5sae_kl_finetune. Python
4interp_tools. Python
3OthelloUnderstanding. Jupyter Notebook
2token-difr. Python
2rl_oocr. Python
1dictionary_learning. Python
1trl_train_demo. Python
1whitebox_evals. Jupyter Notebook
1