Rare find

Native Sparse Attention. A faster way to run AI attention calculations by skipping unnecessary parts.

github.com/lucidrains/native-sparse-attention-pytorch

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

August 2025
  • Release —0.2.3
  • 0.2.3
  • credit assignment
  • Merge pull request #33 from StrongSpoon/mask_fix
July 2025
  • [bugfix] add mask for offs_n
June 2025
  • Release —0.2.2
  • address https://github.com/lucidrains/native-sparse-attention-pytorch…
May 2025
  • Release —0.2.1
  • fix maximum tracking in triton
March 2025
  • turns out gemma v3 decided on qk norm over softcapping
  • start fashioning a version with attention softclamping a la gemma and…
  • Release —0.2.0
  • release new compression block hparams
  • Merge pull request #25 from Mr-Grin/main
  • additional assertion to make sure compress_block_sliding_stride must …
  • minor fix
  • minor modifications to accomodate compress_block_sliding_stride
  • compress_block_sliding_stride
  • Release —0.1.27
  • 0.1.27
  • Merge pull request #24 from Pasewark/token_inference_fix
  • Small change so token embeddings aren't looked up for past tokens dur…
  • Release —0.1.26
  • 0.1.26
  • address https://github.com/lucidrains/native-sparse-attention-pytorch…
  • not using attn biases
  • Release —0.1.25
  • Release —0.1.24
  • Release —0.1.23
  • Release —01.22