This is your work, valued
linear-attention-and-beyond-slides. TeX
120flash-linear-rnn. Implementations of various linear RNN layers using pytorch and triton
55TN-PCFG. source code of NAACL2021 "PCFGs Can Do Better: Inducing Probabilistic Context-Free Grammars with Many Symbols“ and ACL2021 main conference "Neural Bilexicalized PCFG Induction"
52mamba-triton. Python
52fla-tilelang. Python
37pointer-net-for-nested. The official implementation of ACL2022``Bottom-Up Constituency Parsing and Nested Named Entity Recognition with Pointer Networks''
34gated_linear_attention_layer. Python
32span-based-dependency-parsing. Source code of ACL2022 "Headed-Span-Based Projective Dependency Parsing" and "Combining (second-order) graph-based and headed-span-based projective dependency parsing
16second-order-neural-dmv. source code of COLING2020 "Second-Order Unsupervised Neural Dependency Parsing"
16disco-pointer. Official Implementation of ACL2023: Don't Parse, Choose Spans! Continuous and Discontinuous Constituency Parsing via Autoregressive Span Selection
14TN-LCFRS. Official Implementation of ACL2023: Unsupervised Discontinuous Constituency Parsing with Mildly Context-Sensitive Grammars
12hope-fla. Python
7sustcsonglin.github.io. HTML
6path_attention.
6gnn-sdp. Findings of EMNLP'22: Semantic Dependency Parsing with Edge GNNs
3lit-gpt. Hackable implementation of state-of-the-art open-source LLMs based on nanoGPT. Supports flash attention, 4-bit and 8-bit quantization, LoRA and LLaMA-Adapter fine-tuning, pre-training. Apache 2.0-licensed.
3cuda-inside-algs.
3FlagAttention. A collection of memory efficient attention operators implemented in the Triton language.
2triton-inside. Using triton to implement inside algorithms
2unlocking_state_tracking. Expanding linear RNN state-transition matrix eigenvalues to include negatives improves state-tracking tasks and language modeling without added training or inference costs.
2pytorch-custom-mma. C++
2dolomite-engine. Dolomite Engine is a library for pretraining/finetuning LLMs
1nanokitchen. Parallel Associative Scan for Language Models
1mamba.py. An efficient Mamba implementation in PyTorch and MLX.
1streaming-llm. Efficient Streaming Language Models with Attention Sinks
1lcfrs-supertagger. Python
1TinyLlama. The TinyLlama project is an open endeavor to pretrain a 1.1B Llama model on 3 trillion tokens.
1recurrent-chunked-models-regular-languages. Code of "Recurrent Transformers Trade-off Parallelism for Length Generalization on Regular Languages"
1sane_tikz. Reconquer the canvas: beautiful Tikz figures without clunky Tikz code
1shtthesis. An unofficial LaTeX thesis template for ShanghaiTech University.
1stk. Python
1zoology. Understand and test language model architectures on synthetic tasks.
1sustcsonglin_old.github.io. :page_facing_up: Elegant & friendly homepage (bio, tech portfolio, resume, doc...) template with Markdown and VuePress
1safari. Convolutions for Sequence Modeling
1linear-rnn-pytorch.
1cuda-playground. Cuda
1transformers_ssm_copy. Python
1