PSFT. [ICLR 2026] PSFT is a trust-region–inspired fine-tuning objective that views SFT as a policy gradient method with constant advantages, constraining policy drift to stabilize training and improve generalization.
39Hybrid-Policy-Distillation. [ICML 2026] Hybrid Policy Distillation (HPD) is a practical distillation framework for reasoning-oriented language models. This repository contains the code, configurations, and experiment assets for the project.
23adaptive_decoding. [ICML2024]Adaptive decoding balances the diversity and coherence of open-ended text generation.
19weak-to-strong-preference-optimization. [ICLR 2025 Spotlight] Weak-to-strong preference optimization: stealing reward from weak aligned model
18ReAligner. [NeurIPS 2025] A Framework for Flexible Realignment of Language Models during Training and Inference
6penalty_decoding. [EMNLP2023]"Penalty Decoding: Well Suppress the Self-Reinforcement Effect in Open-Ended Text Generation"
2