QVLM. [NeurIPS'24]Efficient and accurate memory saving method towards W4A4 large multi-modal models.
APQ-DM. This is the official pytorch implementation for the paper: Towards Accurate Post-training Quantization for Diffusion Models.(CVPR24 Poster Highlight)