dpo-qwen2-0.5b. 一个可复现的DPO微型项目:基于Qwen2-0.5B-Instruct模型与ultrafeedback_binarized,包含TensorBoard训练曲线及评估/调试脚本。A reproducible DPO (TRL) mini-project: Qwen2-0.5B-Instruct + ultrafeedback_binarized, TensorBoard curves, and evaluation/debug scripts.

github.com/ZHAOoops/dpo-qwen2-0.5b

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.