Rare find

Reading List. A collection of research papers on making AI models faster and smaller.

github.com/evanmiller/LLM-Reading-List

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

December 2023
  • Mamba, PowerInfer
September 2023
  • Go Wider Instead of Deeper
  • Pruning vs Quantization: Which is Better?
August 2023
  • Intriguing Properties of Quantization at Scale
  • Large Transformer Model Inference Optimization (Lilian Weng)
  • The Transformer Family Version 2.0
  • NoPE + ReRoPE
  • Mixture of Experts section
  • Augmenting Self-attention with Persistent Memory
  • Up or Down? Adaptive Rounding for Post-Training Quantization
  • SqueezeLLM: Dense-and-Sparse Quantization
  • QuIP: 2-Bit Quantization of Large Language Models With Guarantees
July 2023
  • initial commit