Rare find

AlignProp. AlignProp uses direct reward backpropogation for the alignment of large-scale text-to-image diffusion models. Our method is 25x more sample and compute efficient than reinforcement learning methods (PPO) for finetuning Stable Diffusion

github.com/mihirp1998/AlignProp

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.