This is your work, valued
Tencent| PhD, Peking University | Multimodal LLMs and Spiking Neural Networks
spikformer. ICLR 2023, Spikformer: When Spiking Neural Network Meets Transformer
★ 407STSGCN_Pytorch. Python
★ 5ZK-Zhou. Config files for my GitHub profile.
★ 2Spike-Driven-Transformer. Spike-Driven Transformer
★ 1Hetero-RL. Group Expectation Policy Optimization for Heterogeneous Reinforcement Learning
★ 171DeepSeek-Math. DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
★ 3.4kWISE. [ICML 2026🔥] WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
★ 212QSD-Transformer. Offical implementation of "Quantized Spike-driven Transformer" (ICLR2025)
★ 34FlashMLA. FlashMLA: Efficient Multi-head Latent Attention Kernels
★ 13kMoH. MoH: Multi-Head Attention as Mixture-of-Head Attention
★ 310SpikeYOLO. Offical implementation of "Integer-Valued Training and Spike-Driven Inference Spiking Neural Network for High-performance and Energy-efficient Object Detection" (ECCV2024 Best Paper Candidate)
★ 251MetaLA. Offical implementation of "MetaLA: Unified Optimal Linear Approximation to Softmax Attention Map" (NeurIPS2024 Oral)
★ 36Spike-Driven-Transformer-V3. Offical implementation of "Scaling Spike-driven Transformer with Efficient Spike Firing Approximation Training" (IEEE T-PAMI2025)
★ 116IC-Light. More relighting!
★ 8.5ktransformer_vq. Official implementation of 'Transformer-VQ: Linear-Time Transformers via Vector Quantization' (ICLR 2024)
★ 199SpikedAttention. This is simple code of SpikedAttention (Neurips 2024)
★ 23LLaVA-CoT. [ICCV 2025] LLaVA-CoT, a visual language model capable of spontaneous, systematic reasoning
★ 2.1kai-for-grant-writing. A curated list of resources for using LLMs to develop more competitive grant applications.
★ 4.2kViewCrafter. [TPAMI 2025] ViewCrafter: Taming Video Diffusion Models for High-fidelity Novel View Synthesis
★ 1.6kAMBER. An LLM-free Multi-dimensional Benchmark for Multi-modal Hallucination Evaluation
★ 173outlier-free-transformers. Python
★ 47AesExpert. [ACMMM 2024] AesExpert: Towards Multi-modality Foundation Model for Image Aesthetics Perception
★ 105AesBench. An expert benchmark aiming to comprehensively evaluate the aesthetic perception capacities of MLLMs.
★ 261Uniaa. Unified Multi-modal IAA Baseline and Benchmark
★ 94mamba. Mamba SSM architecture
★ 19kQKFormer. Offical code of "QKFormer: Hierarchical Spiking Transformer using Q-K Attention"
★ 150MotionInversion. [SIGGRAPH 2025] Official implementation of 'Motion Inversion For Video Customization'
★ 153Spike-Driven-Transformer-V2. Offical implementation of "Spike-driven Transformer V2: Meta Spiking Neural Network Architecture Inspiring the Design of Next-generation Neuromorphic Chips" (ICLR2024)
★ 231Masked-Spiking-Transformer. [ICCV-23 + IJCV-26] Masked Spiking Transformer
★ 32SGLFormer. Spiking Global-Local Fusion Transformer
★ 23SGLFormer. Spiking Global-Local Fusion Transformer
★ 2ProLLaMA. A Protein Large Language Model for Multi-Task Protein Language Processing
★ 207ECDFormer. 【Nature Computational Science 2025🔥】Deep peak property learning for efficient chiral molecules ECD spectra prediction
★ 51mamba-minimal. Simple, minimal implementation of the Mamba SSM in one file of PyTorch.
★ 3kMiniGPT-4. Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)
★ 26kQwen-VL. The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
★ 6.7kPiCO. [ICLR'25] PiCO: Peer Review in LLMs based on the Consistency Optimization, https://arxiv.org/pdf/2402.01830
★ 36Awesome-Spiking-Neural-Networks. A paper list of spiking neural networks, including papers, codes, and related websites. 本仓库收集脉冲神经网络相关的顶会顶刊以及CNS论文和代码,正在持续更新中。
★ 808HARDVS. [AAAI-2024, IJCV-2026] HARDVS: Revisiting Human Activity Recognition with Dynamic Vision Sensors
★ 57MoE-LLaVA. 【TMM 2025🔥】 Mixture-of-Experts for Large Vision-Language Models
★ 2.3kAestheval. Code for the paper "Understanding Aesthetics with Language: A Photo Critique Dataset for Aesthetic Assessment"
★ 103BAID. Python
★ 101Vim. [ICML 2024] Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
★ 3.9kVMamba. VMamba: Visual State Space Models,code is based on mamba
★ 3.2kSSAH-adversarial-attack. Code for the paper "Frequency-driven Imperceptible Adversarial Attack on Semantic Similarity"
★ 62LM4VisualEncoding. [ICLR 2024 (Spotlight)] "Frozen Transformers in Language Models are Effective Visual Encoder Layers"
★ 245xtuner. A Next-Generation Training Engine Built for Ultra-Large MoE Models
★ 5.2kVary. [ECCV 2024] Official code implementation of Vary: Scaling Up the Vision Vocabulary of Large Vision Language Models.
★ 1.9kQ-Bench. An archived version of Q-Bench. We will make updates in https://github.com/q-future/Q-Bench in the future.
★ 12Image-Color-Aesthetics-and-Quality-Assessment. 🔥[ICCV 2023, Official Code] for paper "Thinking Image Color Aesthetics Assessment: Models, Datasets and Benchmarks". Official Weights and Demos provided. 首个面向图像色彩主观美学评估的数据集、算法和benchmark.
★ 214StableLLAVA. Official repo for StableLLAVA
★ 94Otter. 🦦 Otter, a multi-modal model based on OpenFlamingo (open-sourced version of DeepMind's Flamingo), trained on MIMIC-IT and showcasing improved instruction-following and in-context learning ability.
★ 3.4kVideo-LLaVA. 【EMNLP 2024🔥】Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
★ 3.5kImage-Aesthetics-and-Quality-Assessment. 🔥[ACMMM 2023, Official Code] for paper "EAT: An Enhancer for Aesthetics-Oriented Transformers". Official Weights and Demos provided. 目前是地表最强开源美学评估模型之一.
★ 193EMCL. [NeurIPS 2022 Spotlight] Expectation-Maximization Contrastive Learning for Compact Video-and-Language Representations
★ 148GraphMotion. [NeurIPS 2023] Act As You Wish: Fine-Grained Control of Motion Diffusion Model with Hierarchical Semantic Graphs
★ 129MQ-Det. Official PyTorch implementation of "Multi-modal Queried Object Detection in the Wild" (accepted by NeurIPS 2023)
★ 346Dynamic-Token-Pruning. Official Pytorch implementation of Dynamic-Token-Pruning (ICCV2023)
★ 23DynamicViT. [NeurIPS 2021] [T-PAMI] DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
★ 668LLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kMedusa. Medusa: Simple Framework for Accelerating LLM Generation with Multiple Decoding Heads
★ 2.8kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kbaby-llama2-chinese. 用于从头预训练+SFT一个小参数量的中文LLaMa2的仓库;24G单卡即可运行得到一个具备简单中文问答能力的chat-llama2.
★ 2.9kloss-landscape. Code for visualizing the loss landscape of neural nets
★ 3.2kEmbodiedGPT_Pytorch. Python
★ 348MADAv2. MADAv2: Advanced Multi-Anchor Based Active Domain Adaptation Segmentation
★ 25Spike-Driven-Transformer. Offical implementation of "Spike-driven Transformer" (NeurIPS2023)
★ 317Fast-SNN. Python
★ 73RRHF. [NIPS2023] RRHF & Wombat
★ 805ChatLaw. ChatLaw:A Powerful LLM Tailored for Chinese Legal. 中文法律大模型
★ 7.6kAwesome-LLM. Awesome-LLM: a curated list of Large Language Model
★ 27kNeuroCLIP. Python
★ 9MAE-Lite. Official implement for ICML2023 paper: "A Closer Look at Self-Supervised Lightweight Vision Transformers"
★ 151ConvNeXt-V2. Code release for ConvNeXt V2 model
★ 2.1kVanillaNet. Python
★ 821SAVC. [CVPR 2023] Learning with Fantasy: Semantic-Aware Virtual Contrastive Constraint for Few-Shot Class-Incremental Learning
★ 75qlora. QLoRA: Efficient Finetuning of Quantized LLMs
★ 11kPointGPT. [NeurIPS 2023] PointGPT: Auto-regressively Generative Pre-training from Point Clouds
★ 245Spikingformer. Spikingformer: A Key Foundation Model for Spiking Neural Networks (AAAI 2026)
★ 149Spikingformer-CML. Enhancing the Performance of Transformer-based Spiking Neural Networks by SNN-optimized Downsampling with Precise Gradient Backpropagation
★ 49Swin-MAE. Pytorch implementation of Swin MAE https://arxiv.org/abs/2212.13805
★ 107Compact-Transformers. Escaping the Big Data Paradigm with Compact Transformers, 2021 (Train your Vision Transformers in 30 mins on CIFAR-10 with a single GPU!)
★ 546LLM-in-Vision. Recent LLM-based CV and related works. Welcome to comment/contribute!
★ 871FCViT. A Close Look at Spatial Modeling: From Attention to Convolution
★ 92hivit. Jupyter Notebook
★ 76Parallel-Spiking-Neuron. Python
★ 56SNNCKA. Pytorch implementation of ANN-SNN representation similarity analysis (TMLR 2023)
★ 13ml-cvnets. CVNets: A library for training computer vision networks
★ 2kSpike-Element-Wise-ResNet. Deep Residual Learning in Spiking Neural Networks
★ 195efficientvit. Efficient vision foundation models for high-resolution generation and perception.
★ 3.3ksegment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kRWKV-LM. RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.
★ 15kEfficient-Computing. Efficient computing methods developed by Huawei Noah's Ark Lab
★ 1.3kConvMAE. ConvMAE: Masked Convolution Meets Masked Autoencoders
★ 531SparK. [ICLR'23 Spotlight🔥] The first successful BERT/MAE-style pretraining on any convolutional network; Pytorch impl. of "Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling"
★ 1.4kspikformer. ICLR 2023, Spikformer: When Spiking Neural Network Meets Transformer
★ 407CAE. This is a PyTorch implementation of “Context AutoEncoder for Self-Supervised Representation Learning"
★ 199detectron2. Detectron2 is a platform for object detection, segmentation and other visual recognition tasks.
★ 35kMIMDet. [ICCV 2023] You Only Look at One Partial Sequence
★ 343Attention-driven-Masking-and-Throwing. Python
★ 73SNN-Neural-Similarity-Static. Deep Spiking Neural Networks with High Representation Similarity Model Visual Pathways of Macaque and Mouse
★ 18mae. PyTorch implementation of MAE https//arxiv.org/abs/2111.06377
★ 8.4ksyops-counter. SyOPs counter for spiking neural networks
★ 78Q-ViT. The official implementation of the NeurIPS 2022 paper Q-ViT.
★ 106rebiber. A simple tool to update bib entries with their official information (e.g., DBLP or the ACL anthology).
★ 3kAwesome-Transformer-Attention. An ultimately comprehensive paper list of Vision Transformer/Attention, including papers, codes, and related websites
★ 5.1kMAE-pytorch. Unofficial PyTorch implementation of Masked Autoencoders Are Scalable Vision Learners
★ 2.7kPoint-MAE. [ECCV2022] Masked Autoencoders for Point Cloud Self-supervised Learning
★ 638spikingjelly. SpikingJelly is an open-source deep learning framework for Spiking Neural Network (SNN) based on PyTorch.
★ 2.1kSwin-Transformer. This is an official implementation for "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows".
★ 16kfocal_loss_pytorch. A PyTorch Implementation of Focal Loss.
★ 991auto-feature-extraction-method-for-e-tongue. Python
★ 5slam_gmapping. http://www.ros.org/wiki/slam_gmapping
★ 734