This is your work, valued
Unsupervised-Elicitation. Python
41MisleadLM. Official Code for our paper: "Language Models Learn to Mislead Humans via RLHF""
20GDsuite. A toy eval suite for tracing generalization dynamics of LM pre-training
19storyboard.
1Introduction-to-Machine-Learning. Python
1mend. MEND: Fast Model Editing at Scale
1