ArtQuantization. ArtQuantization is developed for quantizing Large Language Models, focusing on optimizing the memory usage and performance. This repository provides experimental results of quantizing models such as Qwen2.5 using different algorithms like AWQ and GPTQ, and demonstrates the memory requirements under various graphics card configurations.

github.com/Artessay/ArtQuantization

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.