sycophancy-eval. datasets from the paper "Towards Understanding Sycophancy in Language Models"
activation-steering. finding a bias-correlated activation vector for llama-2-7b-chat
trlx. A repo for distributed training of language models with Reinforcement Learning via Human Feedback (RLHF)