A_Dynamic_Multi-Modal_Deep_Reinforcement_Learning_Framework_for_3D_Bin_Packing_Problem. the Pytorch implementation of A Dynamic Multi-Modal Deep Reinforcement Learning Framework for 3D Bin Packing Problem
11On_Policy_Distillation_Paper_List.
82D-Coordinate-System-for-ICL. [EMNLP 2024 Main] Official implementation of the paper "Unveiling In-Context Learning: A Coordinate System to Understand Its Working Mechanism".
6Awesome_SFT-RLVR_Mechanism. A curated collection of papers on the roles and mechanisms of SFT and RLVR in LLM reasoning training.
6On_Policy_SFT. On-Policy SFT for Adaptive Reasoning (Reproducible Implementation)
2