PC-DARTS. PC-DARTS:Partial Channel Connections for Memory-Efficient Differentiable Architecture Search
442qa-lora. Official PyTorch implementation of QA-LoRA
147BNET. Batch Normalization with Enhanced Linear Transformation
53Trained-Rank-Pruning. Pytorch implementation of TRP
44WLQ. caffe implementation of single level quantization
19gptqlora. GPTQLoRA: Efficient Finetuning of Quantized LLMs with GPTQ
2GPTQ-for-LLaMa. 4 bits quantization of LLaMA using GPTQ, zeros fp16
1diracnets. Training Very Deep Neural Networks Without Skip-Connections
1channel-pruning. Channel Pruning for Accelerating Very Deep Neural Networks
1attention-transfer. Improving Convolutional Networks via Attention Transfer (ICLR 2017)
13D-Machine-Learning. A learning resource repository for 3D machine learning
1KVCache-Factory. Unified KV Cache Compression Methods for Auto-Regressive Models
1kvpress. LLM KV cache compression made easy
1img_classification_pk_pytorch. Quickly comparing your image classification models with the state-of-the-art models (such as DenseNet, ResNet, ...)
1