Rare find

Llama-AVSR. Official Pytorch implementation of "Large Language Models are Strong Audio-Visual Speech Recognition Learners" [ICASSP 2025] and "Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs" [ICASSP 2026].

github.com/umbertocappellazzo/Llama-AVSR

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.