gated_attention. The official implementation for [NeurIPS2025 Oral] Gated Attention for Large Language Models: Non-linearity, Sparsity, and Attention-Sink-Free
973EMoE. Official PyTorch Implementation of EMoE: Unlocking Emergent Modularity in Large Language Models [main conference @ NAACL2024]
39RMoE. Official implementation of RMoE (Layerwise Recurrent Router for Mixture-of-Experts)
33EEG-Cross-Subject-Emotion-Recognition. Python
10HMA. HMA: Heterogenous Memory Augmented Neural Networks
5Tuning-keys-v.s.-values. Official PyTorch Implementation of Empirical Study on Updating Key-Value Memories in Transformer Feed-forward Layers [Tiny Paper @ ICLR 2024]
4DL_project. Deep learning class project -- A rational search engine
3eegnet_pytorch. EEGNet implementation in PyTorch
1MemLMM. An implementation of model parallel autoregressive transformers on GPUs, based on the DeepSpeed library.
1