Hao Sun

Expert
@holarissun

PhD in Reinforcement Learning, LLM Alignment, RLHF

PanelGPT. We introduce new zero-shot prompting magic words that improves the reasoning ability of language models: panel discussion!

174

RewardModelingBeyondBradleyTerry. official implementation of ICLR'2025 paper: Rethinking Bradley-Terry Models in Preference-based Reward Modeling: Foundations, Theory, and Alternatives

73

Prompt-OIRL. code for paper Query-Dependent Prompt Evaluation and Optimization with Offline Inverse Reinforcement Learning

45

RewardShifting. Code for NeurIPS 2022 paper Exploiting Reward Shifting in Value-Based Deep RL

29

embedding-based-llm-alignment. Codebase for Paper Reusing Embeddings: Reproducible Reward Model Research in Large Language Model Alignment without GPUs

22

PCHID_code. Code for [NeurIPS'2019 Spotlight] Policy Continuation with Hindsight Inverse Dynamics

15

InverseRLmeetsLLMs.

10

DOMIAS. Python

6

Accountable-Offline-RL. Code for NeurIPS 2023 paper Accountability in Offline Reinforcement Learning: Explaining Decisions with a Corpus of Examples

5

HandsOnTransformers. attention is all you need!

5

Inverse-RLignment. inverse reinforcement learning for LLM alignment

4

DAUC. Code for Latent Density Models for Uncertainty Categorization

3

Exploiting-Exploitation. Python

2

MOPA. Jupyter Notebook

2

NPSCO. Code for Novel Policy Seeking with Constrained Optimization

2

Action-Refined-Temporal-Difference. Jupyter Notebook

2

Prompt-Engineering-Guide. 🐙 Guides, papers, lecture, notebooks and resources for prompt engineering

1

LeetCodeSolution. logs for my leetcoding fall 2023

1

BenchmarkPromptsWithResponses. Every prompt engineering paper should provide not only on-average performance of the prompting strategy, but should also release the responses to facilitate future research and avoid repeatedly calling the LLMs for the same queries+prompts.

1

Data-Centric-OPE. Python

1

Policy-Continuation-with-Hindsight-Inverse-Dynamics. Jupyter Notebook

1

SqueezingOpenReview. Python

1
22
Apply