Postdoctoral researcher - Stanford
minimal-LRU. Non official implementation of the Linear Recurrent Unit (LRU, Orvieto et al. 2023)
63Online-learning-LR-dependencies. Implementation of the "Online learning of long-range dependencies" paper, NeurIPS 2023
21Forward-propagation-errors-through-time. Jupyter Notebook
4The-emergence-of-sparse-attention. Companion codebase for the paper "The emergence of sparse attention: impact of data distribution and benefits of repetition"
2Vanishing-and-exploding-gradients-are-not-the-end-of-the-story. Code for the paper "Recurrent neural networks: vanishing and exploding gradients are not the end of the story".
1