This is your work, valued
awesome-RLAIF. A continually updated list of literature on Reinforcement Learning from AI Feedback (RLAIF)
205internally-rewarded-rl. [ICML 2023] Code for paper "Internally Rewarded Reinforcement Learning"
13vanilla-RLAIF-pipeline. An implementation of a vanilla RLAIF pipeline, utilizing GPT-2-Large for the summarization task with the TL;DR dataset.
8robotic-occlusion-reasoning. PyTorch implementation of "Robotic Occlusion Reasoning for Efficient Object Existence Prediction" (IROS 2021)
8UNO-wiki. 记录UNO自定义版本规则。
1