Wei Xiong

Expert
@WeiXiongUST

Ph.D. Student in computer science at UIUC; machine learning theory and RLHF.

Building-Math-Agents-with-Multi-Turn-Iterative-Preference-Learning. This is an official implementation of the paper ``Building Math Agents with Multi-Turn Iterative Preference Learning'' with multi-turn DPO and KTO.

32

Decentralized-Proximal-Algorithm-with-Variance-Reduction. This is the code used for the paper "PMGT-VR: A decentralized proximal-gradient algorithmic framework with variance reduction", prepint.

16

multi-armed-bandit-test-framework. This is the code about multi_armed bandit used for my undergraduate thesis.

5

LMFlow_RAFT_Dev. This is a sub-branch for developing RAFT algorithm.

4

MATH6913U.

2

CS598. Python

2

Online-RLHF. Python

2

MPMAB_BEACON. This is the official implementation for the paper "Heterogeneous Multi-player Multi-armed Bandits: Closing the Gap and Generalization" in NeurIPS 2021.

2

feedbackagent. Python

2

dpo_math. Python

1

vllm_eval. Python

1

Observe_then_Incentivize. This is the official implementation for the paper "(Almost) Free Incentivized Exploration from Decentralized Learning Agents" in NeurIPS 2021.

1

multi_player_multi_armed_bandit_algorithms. Implementation of state-of-the-art multi-player multi-armed bandit problem algorithms.

1

RLHF-Reward-Modeling-dev. Recipes to train reward model for RLHF.

1

ToRA. ToRA is a series of Tool-integrated Reasoning LLM Agents designed to solve challenging mathematical reasoning problems by interacting with tools [ICLR'24].

1

code_trn. Python

1

iterative-rlhf. Python

1

infer_math. Python

1
18
Apply