Hi, I'm Kai Yang. I earned my master's degree from Tsinghua in 2025, specializing in RL and LLM. I am presently employed at Tencent Hunyuan.
d3po. [CVPR 2024] Code for the paper "Using Human Feedback to Fine-tune Diffusion Models without Any Reward Model"
244TaskAllocation. [EAAI] A two-stage reinforcement learning-based approach for multi-entity task allocation.
39Graduation-project-design. 采用模拟退火策略优化的免疫算法 解决无人机协同分配问题
24DRND. [ICML 2024]Exploration and Anti-exploration with Distributional Random Network Distillation
18EntroPIC. EntroPIC: Towards Stable Long-Term Training of LLMs via Entropy Stabilization with Proportional-Integral Control
13MachineLearning. 吴恩达机器学习作业代码
3yk7333. Config files for my GitHub profile.
2DIP. 数字图像处理
1