Rare find

Vision Transformer. AI model that learns to recognize images by looking at them piece by piece, like a human reading a page.

github.com/lucidrains/vit-pytorch

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

June 2026
  • vibe out a vit with sparse gating with a thresholded l1 loss for pers…
  • credit assignment
  • 1.24.0 - releasing vit-5 contributed by @pranoyr
  • Adding vit-5 (#366)
  • always attention gate
  • allow for sigreg of the slots, for some part-whole ssl losses
  • credit assignment
  • final thought for the wwt, multiple hierarchices of slots should be a…
  • kick it up a notch for wwt, allow for multiple hierarchies of slots, …
  • 1.23.4
  • added the supernetwork for two stage distillation (#365)
  • add the autoencoding head, which may matter
  • allow for an l1 norm after the slot attention in wwt, add register to…
  • vibe in a copy of the what-where transformer proposed by Yoshihashi e…
  • release jet vit contributed by @pranoyr
  • added jet_vit (#364)
May 2026
  • throw out something to be explored
  • add ability to regularize subspaces for the vit with decorr loss
  • have gemini work out sequential w/ cache and parallel parity for moss…
  • add moss to accept video wrapper, aimed towards SRT-H
February 2026
  • Release —1.17.8
  • Release —1.17.7
January 2026
  • Release —1.17.6
  • Release —1.17.5
  • Release —1.17.4
  • Release —1.17.3
  • Release —1.17.2
December 2025
  • Release —1.17.1
  • Release —1.16.5
  • Release —1.16.4