This is your work, valued
UMoE. Python
23XMoE. Python
15TopicAns. Implementation for TopicAns
4MixEncoder. Python
2fairseq_moe. Python
1flash-linear-attention. 🚀 Efficient implementations of state-of-the-art linear attention models in Torch and Triton
1news-please. news-please - an integrated web crawler and information extractor for news that just works
1MoE-Megatron-DeepSpeed. Ongoing research training transformer language models at scale, including: BERT & GPT-2
1RetroMAE. Codebase for RetroMAE and beyond.
1hf_translation.
1ysngki.github.io. HTML
1