This is your work, valued
Ph.D. Student in computer science at UIUC; machine learning theory and RLHF.
Building-Math-Agents-with-Multi-Turn-Iterative-Preference-Learning. This is an official implementation of the paper ``Building Math Agents with Multi-Turn Iterative Preference Learning'' with multi-turn DPO and KTO.
32Decentralized-Proximal-Algorithm-with-Variance-Reduction. This is the code used for the paper "PMGT-VR: A decentralized proximal-gradient algorithmic framework with variance reduction", prepint.
16multi-armed-bandit-test-framework. This is the code about multi_armed bandit used for my undergraduate thesis.
5LMFlow_RAFT_Dev. This is a sub-branch for developing RAFT algorithm.
4MATH6913U.
2CS598. Python
2Online-RLHF. Python
2MPMAB_BEACON. This is the official implementation for the paper "Heterogeneous Multi-player Multi-armed Bandits: Closing the Gap and Generalization" in NeurIPS 2021.
2feedbackagent. Python
2dpo_math. Python
1vllm_eval. Python
1Observe_then_Incentivize. This is the official implementation for the paper "(Almost) Free Incentivized Exploration from Decentralized Learning Agents" in NeurIPS 2021.
1multi_player_multi_armed_bandit_algorithms. Implementation of state-of-the-art multi-player multi-armed bandit problem algorithms.
1RLHF-Reward-Modeling-dev. Recipes to train reward model for RLHF.
1ToRA. ToRA is a series of Tool-integrated Reasoning LLM Agents designed to solve challenging mathematical reasoning problems by interacting with tools [ICLR'24].
1code_trn. Python
1iterative-rlhf. Python
1infer_math. Python
1