Research Engineer @ NatWest Group. Previously: RA @ Imperial College London. Advancing multimodal AI.
Llama-AVSR. Official Pytorch implementation of "Large Language Models are Strong Audio-Visual Speech Recognition Learners" [ICASSP 2025] and "Mitigating Attention Sinks and Massive Activations in Audio-Visual Speech Recognition with LLMs" [ICASSP 2026].
64PETL_AST. This is the official repository of the papers "Parameter-Efficient Transfer Learning of Audio Spectrogram Transformers" [IEEE MLSP 2024] and "Efficient Fine-tuning of Audio Spectrogram Transformers via Soft Mixture of Adapters" [Interspeech 2024].
41Omni-AVSR. Official Pytorch implementation of "Omni-AVSR: Towards Unified Multimodal Speech Recognition with Large Language Models" [IEEE ICASSP 2026].
38CL_Anthology. An anthology of recent continual learning papers, where people interested in this fascinating topic can start discovering its multidimensional representations.
10CL_SLU. "An Investigation of the Combination of Rehearsal and Knowledge Distillation in Continual Learning for Spoken Language Understanding", accepted at INTERSPEECH 2023.
8Dr-SHAP-AV. Official Pytorch implementation of "Dr. SHAP-AV: Decoding Relative Modality Contributions via Shapley Attribution in Audio-Visual Speech Recognition" [Interspeech 2026, Long Track].
7umbertocappellazzo.github.io. HTML
1