PanelGPT. We introduce new zero-shot prompting magic words that improves the reasoning ability of language models: panel discussion!
174RewardModelingBeyondBradleyTerry. official implementation of ICLR'2025 paper: Rethinking Bradley-Terry Models in Preference-based Reward Modeling: Foundations, Theory, and Alternatives
73Prompt-OIRL. code for paper Query-Dependent Prompt Evaluation and Optimization with Offline Inverse Reinforcement Learning
45RewardShifting. Code for NeurIPS 2022 paper Exploiting Reward Shifting in Value-Based Deep RL
29embedding-based-llm-alignment. Codebase for Paper Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs
22PCHID_code. Code for [NeurIPS'2019 Spotlight] Policy Continuation with Hindsight Inverse Dynamics
15InverseRLmeetsLLMs.
10DOMIAS. Python
6Accountable-Offline-RL. Code for NeurIPS 2023 paper Accountability in Offline Reinforcement Learning: Explaining Decisions with a Corpus of Examples
5HandsOnTransformers. attention is all you need!
5Inverse-RLignment. inverse reinforcement learning for LLM alignment
4DAUC. Code for Latent Density Models for Uncertainty Categorization
3Exploiting-Exploitation. Python
2MOPA. Jupyter Notebook
2NPSCO. Code for Novel Policy Seeking with Constrained Optimization
2Action-Refined-Temporal-Difference. Jupyter Notebook
2Prompt-Engineering-Guide. 🐙 Guides, papers, lecture, notebooks and resources for prompt engineering
1LeetCodeSolution. logs for my leetcoding fall 2023
1BenchmarkPromptsWithResponses. Every prompt engineering paper should provide not only on-average performance of the prompting strategy, but should also release the responses to facilitate future research and avoid repeatedly calling the LLMs for the same queries+prompts.
1Data-Centric-OPE. Python
1Policy-Continuation-with-Hindsight-Inverse-Dynamics. Jupyter Notebook
1SqueezingOpenReview. Python
1