This is your work, valued
Ph.D. student at The Chinese University of Hong Kong, Shenzhen.
ReMax. Code for Paper (ReMax: A Simple, Efficient and Effective Reinforcement Learning Method for Aligning Large Language Models)
202GEM. Code for Paper (Preserving Diversity in Supervised Fine-tuning of Large Language Models)
58policy_optimization. Code for Paper (Policy Optimization in RLHF: The Impact of Out-of-preference Data)
29cold_start_rl. Code for Blog Post: Can Better Cold-Start Strategies Improve RL Training for LLMs?
20KnapsackRL. Python
19RL-PPO-Keras. Proximal Policy Optimization(PPO) with Keras Implementation
17HyperDQN. Code for ICLR 2022 Paper (HyperDQN: A Randomized Exploration Method for Deep Reinforcement Learning)
12Face-Recognition-TX2. Face Recognition on NVIDIA TX2
9verl. veRL: Volcano Engine Reinforcement Learning for LLM
8ISWBC. Code for NeurIPS 2023 Paper (Imitation Learning from Imperfection: Theoretical Justifications and Algorithms)
8RLX. RLX is an RL codebase based on TensorFlow. It implements algorithms like SAC, ACER, GAIL and TRPO. It is easy to use.
3ILwSD. Python
3CVPR-XJTU-2018. 2018 Spring Course, Computer Vision and Pattern Recognition, in XJTU
3offline_rl. Python
2Machine-Learning-XJTU-2018. The course of machine learning I take in XJTU, 2018 spring, guided by Prof Deyu Meng
2High-Dimension-Data-Analysis. High Dimensional Data Analysis: Lasso, Compressed Sensing, AdaBoost
1bib-merge. Python
1Machine-Learning. This is my implementation of algorithms in the book, 《Machine Learning》, by Zhihua Zhou
1Web_Crawler. Web crawler for several Chinese web pages
1