This is your work, valued

China

Runze Liu

Advanced
@RyanLiu112

Incoming PhD @ HKUST & Master's student @ THU

compute-optimal-tts. Official codebase for "Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling".

288

Awesome-Process-Reward-Models. A comprehensive collection of process reward models.

176

GenPRM. [AAAI 2026] Official codebase for "GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning".

102

MRN. [NeurIPS 2022] Official codebase for "Meta-Reward-Net: Implicitly Differentiable Reward Learning for Preference-based Reinforcement Learning".

26

AttnRL. [ICLR 2026] Official codebase for "Attention as a Compass: Efficient Exploration for Process-Supervised RL in Reasoning Models"

14

RL_parking. HTML

14

genrm-critiques. GenRM-CoT: Data release for verification rationales

2

OpenRLHF-fork. An Easy-to-use, Scalable and High-performance RLHF Framework (70B+ PPO Full Tuning & Iterative DPO & LoRA & Mixtral)

2

RLHF-Reward-Modeling. Recipes to train reward model for RLHF.

2

dpss-exp3-VC-BNF. Voice Conversion Experiments for THUHCSI Course : <Digital Processing of Speech Signals>

2

Awesome-RL-Reasoning-Recipes. Awesome RL Reasoning Recipes ("Triple R")

2

epic. Implements the Equivalent-Policy Invariant Comparison (EPIC) distance for reward functions.

1

evaluating-rewards. Library to compare and evaluate reward functions

1

CLIP4Clip. An official implementation for "CLIP4Clip: An Empirical Study of CLIP for End to End Video Clip Retrieval"

1

LLaMA-VID. Official Implementation for LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

1

factor-world. Decomposing the Generalization Gap in Imitation Learning for Visual Robotic Manipulation (2023)

1

VTF_PAR. [CVPR-2023 Workshop@NFVLR] Official PyTorch implementation of Learning CLIP Guided Visual-Text Fusion Transformer for Video-based Pedestrian Attribute Recognition

1

eai-vc. The repository for the largest and most comprehensive empirical study of visual foundation models for Embodied AI (EAI).

1

rl-teacher-tf. Open source implementation of "Deep Reinforcement Learning from Human Preferences", updating with evolutionary strategies and augmented morphologies from "Reinforcement Learning for Improved Agent Design"

1

LLaVA. Large Language-and-Vision Assistant built towards multimodal GPT-4 level capabilities.

1

diffuser. Code for the paper "Planning with Diffusion for Flexible Behavior Synthesis"

1

mPLUG-Owl. mPLUG-Owl🦉: Modularization Empowers Large Language Models with Multimodality

1

PyRep. A toolkit for robot learning research.

1

rl-teacher-pytorch. Python

1

VideoChat. Python

1

Video-LLaMA. Video-LLaMA: An Instruction-Finetuned Visual Language Model for Video Understanding

1

metaworld. Collections of robotics environments geared towards benchmarking multi-task and meta reinforcement learning

1

RLBench. A large-scale benchmark and learning environment.

1

open_flamingo. An open-source framework for training large multimodal models.

1