This is your work, valued
Awesome-Mamba-Papers. Awesome Papers related to Mamba.
★ 1.4kAwesome-Agent-Memory-Papers. Awesome Papers related to Agent Memory: methods, benchmarks and surveys. Website: https://yyyujintang.github.io/Awesome-Agent-Memory-Papers/
★ 224VMRNN-PyTorch. Official repository for CVPR24 Precognition Workshop Paper: VMRNN: Integrating Vision Mamba and LSTM for Efficient and Accurate Spatiotemporal Forecasting.
★ 165PredFormer. [TMLR2026] Video Prediction Transformers without Recurrence or Convolution
★ 157PostRainBench. Official PyTorch Code for ICLR24 Workshop Paper: PostRainBench
★ 29DMV-Bench. Python
★ 2jenny_blog. Jenny's blog.
★ 1Awesome-Agentic-Time-Series. The paper list of "The Landscape of Agentic Time Series Systems: Architectures, Reliability, and Frontiers."
★ 146AMA-Bench. [ICML 26] An evaluation framework assessing long-context retention and long-horizon memory performance for agentic applications (AMA-bench).
★ 65mem0. Universal memory layer for AI Agents
★ 62kclaw-code. An agent-managed museum exhibit, built in Rust with Gajae-Code / LazyCodex — developed and maintained with no human intervention.
★ 195kHippoRAG. [NeurIPS'24] HippoRAG is a novel RAG framework inspired by human long-term memory that enables LLMs to continuously integrate knowledge across external documents. RAG + Knowledge Graphs + Personalized PageRank.
★ 3.9kPlowPilot. TypeScript
★ 4PlugMem. ICML 2026 · Plug-and-play long-term memory for LLM agents
★ 167alma. ALMA (Automated meta-Learning of Memory designs for Agentic systems) is a framework that meta-learns memory designs to replace human-engineered designs for agentic system.
★ 250Awesome-Agent-Memory. [Up-To-Date] [TMLR 2026] Awesome Agent Memory Paper Resource
★ 175MemGUI-Bench. [ACM MM 2026] MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
★ 47reasoning-bank-slm. An experiment that applies Google Research's `ReasoningBank` technique to Small Language Models. This experiment hopes to show that the same gains from the ReasoningBank paper also applies to much smaller, less capable models.
★ 108Synapse. [ICLR 2024] Trajectory-as-Exemplar Prompting with Memory for Computer Control
★ 70agent-workflow-memory. AWM: Agent Workflow Memory
★ 450alfworld. ALFWorld: Aligning Text and Embodied Environments for Interactive Learning
★ 815webarena. Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents"
★ 1.6kMemSkill. MemSkill: Learning and Evolving Memory Skills for Self-Evolving Agents
★ 551MPR. The official repository for paper "Explicit v.s. Implicit Memory: Exploring Multi-hop Complex Reasoning Over Personalized Information"
★ 1MemoryAgentBench. Open source code for ICLR 2026 Paper: Evaluating Memory in LLM Agents via Incremental Multi-Turn Interactions
★ 413visualwebarena. VisualWebArena is a benchmark for multimodal agents.
★ 484ai-multi-agent-presentation-builder. Python
★ 50LLM_Agent_Memory_Survey.
★ 504Mirage. [CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
★ 294Huaman-Agent-Memory.
★ 112WebShop. [NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
★ 575MemVerse. MemVerse: Multimodal Memory for Lifelong Learning Agents
★ 154sft-qwen2.5-omni-thinker. verl: Volcano Engine Reinforcement Learning for LLMs
★ 42M3-Agent-Training. Python
★ 30ReMe. ReMe: Memory Management Kit for Agents - Remember Me, Refine Me.
★ 3.2kAgent-Memory-Paper-List. The paper list of "Memory in the Age of AI Agents: A Survey"
★ 2.3kVAGEN. World model reasoning RL for multi-turn VLM agents
★ 489A-mem. A-MEM: Agentic Memory for LLM Agents
★ 1.1kA-mem. The code for NeurIPS 2025 paper "A-Mem: Agentic Memory for LLM Agents"
★ 932MemGen. MemGen: Weaving Generative Latent Memory for Self-Evolving Agents
★ 407m3-agent. Python
★ 1.4kNeurIPS24-Optimus-1. [NeurIPS 2024] Official Implementation for Optimus-1: Hybrid Multimodal Memory Empowered Agents Excel in Long-Horizon Tasks
★ 102MWT. Official Repository of End-to-End Implicit Neural Representations for Classification (CVPR 2025)
★ 14functa. Python
★ 166LinearDiff. Python
★ 10test-time-registers. [NeurIPS '25 Spotlight] Official Pytorch implementation of "Vision Transformers Don't Need Trained Registers"
★ 184Vit-RGTS. Open source implementation of "Vision Transformers Need Registers"
★ 218abxlab. [ICLR'26] A Framework for Studying AI Agent Behavior: Evidence from Consumer Choice Experiments
★ 9DeepTutor. DeepTutor: Lifelong Personalized Tutoring. https://deeptutor.info/.
★ 31kdeit. Official DeiT repository
★ 4.4kLLM-Discrete-Tokenization-Survey. [TPAMI 2026] Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey.
★ 83Awesome-Token-Compress. A paper list of some recent works about Token Compress for Vit and VLM
★ 944Show-o. [ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.
★ 2kReAct. [ICLR 2023] ReAct: Synergizing Reasoning and Acting in Language Models
★ 4.1kPH-Reg. Official code for "Vision Transformers with Self-Distilled Registers" (NeurIPS 2025 Spotlight)
★ 35DAP. Official implementation of "Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation".
★ 348ShowUI. [CVPR 2025] Open-source, End-to-end, Vision-Language-Action model for GUI Agent & Computer Use.
★ 1.9kOpen-AutoGLM. An Open Phone Agent Model & Framework. Unlocking the AI Phone for Everyone
★ 26knepa. PyTorch implementation of NEPA
★ 339UnSAM. [NeurIPS 2024] Code release for "Segment Anything without Supervision"
★ 503open_clip. An open source implementation of CLIP.
★ 14kCoVT. [ECCV 2026] Official repo of "Chain-of-Visual-Thought: Teaching VLMs to See and Think Better with Continuous Visual Tokens"
★ 386reconstruction-alignment. [ICLR 2026] Official repo of paper "Reconstruction Alignment Improves Unified Multimodal Models". Unlocking the Massive Zero-shot Potential in Unified Multimodal Models through Self-supervised Learning.
★ 411GUI-Agents-Paper-List. Awesome GUI Agent Paper List
★ 866Awesome-GUI-Agent. 💻 A curated list of papers and resources for multi-modal Graphical User Interface (GUI) agents.
★ 1.2kAwesome-3DGS-Applications. 【TPAMI 2026】A Survey on 3D Gaussian Splatting Applications: Segmentation, Editing, and Generation
★ 393GST. Official code of ``2D Gaussians Spatial Transport for Point-supervised Density Regression'' (GST), AAAI 26.
★ 22GPSToken. The official code for paper "GPSToken: Gaussian Parameterized Spatially-adaptive Tokenization for Image Representation and Generation"
★ 58VARC. Python
★ 251dinov3. Reference PyTorch implementation and models for DINOv3
★ 11klejepa. Python
★ 1.3kdHT. Differentiable Hierarchical Visual Tokenization
★ 45DeepOCR. A reproduction of the Deepseek-OCR model including training
★ 208VLMEvalKit. Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3kFOCUS. [ICLR 2026] FOCUS: Efficient Keyframe Selection for Long Video Understanding
★ 74Spatial-SSRL. [CVPR 2026] Official release of "Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement Learning"
★ 133Awesome-World-Models. A Curated List of Awesome Works in World Modeling, Aiming to Serve as a One-stop Resource for Researchers, Practitioners, and Enthusiasts Interested in World Modeling.
★ 3.2kVisionSelector. VisionSelector: End-to-End Learnable Visual Token Compression for Efficient Multimodal LLMs
★ 65lmms-engine. A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
★ 814Qwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kDeepSeek-OCR. Contexts Optical Compression
★ 24kLLaVA-OneVision-2. Fully Open Framework for Democratized Multimodal Training
★ 1.2kMMVP. Python
★ 365NEO. NEO Series: Native Vision-Language Models from First Principles
★ 878streaming-vlm. StreamingVLM: Real-Time Understanding for Infinite Video Streams
★ 1.1klm-evaluation-harness. A framework for few-shot evaluation of language models.
★ 13klmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kInternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10kLLaVA-NeXT. Python
★ 4.7kcross-modal-information-flow-in-MLLM. This is the official repository for paper: cross-modal information flow in multimodal large language models
★ 44LongLive. Long Video Gen Infrastructure
★ 2.5kLlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kKeye. Python
★ 809MiMo-VL. MiMo-VL
★ 643GLM-V. GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
★ 2.4kVideoLLaMA3. Frontier Multimodal Foundation Models for Image and Video Understanding
★ 1.2kLayerCraft. Official Repo for LayerCraft
★ 18lectures. Python
★ 3.6kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kLLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kLucidFlux. LucidFlux: Caption-Free Photo-Realistic Image Restoration via a Large-Scale Diffusion Transformer, ICLR 2026
★ 1.3kRF-Solver-Edit. [🚀ICML 2025] "Taming Rectified Flow for Inversion and Editing" Using FLUX and HunyuanVideo for image and video editing!
★ 638KV-Edit. [ICCV 2025] Official implementation for KV-Edit: Training-Free Image Editing for Precise Background Preservation
★ 388Long-RL. Long-RL: Scaling RL to Long Sequences (NeurIPS 2025)
★ 727A-Comprehensive-Survey-For-Long-Context-Language-Modeling. A Comprehensive Survey on Long Context Language Modeling
★ 252Q-LLM. This is the official repo of "QuickLLaMA: Query-aware Inference Acceleration for Large Language Models"
★ 54needle-in-a-haystack. Doing simple retrieval from LLM models at various context lengths to measure accuracy
★ 2.4kLLMxMapReduce. Python
★ 873Awesome-LLM-Long-Context-Modeling. 📰 Must-read papers and blogs on LLM based Long Context Modeling 🔥
★ 2.1kRSCP2GAN. Python
★ 21slurm-tutorial. SLURM Tutorial
★ 28Empirical-Study-of-GPT-4o-Image-Gen. An Empirical Study of GPT-4o Image Generation Capabilities
★ 29Janus. Janus-Series: Unified Multimodal Understanding and Generation Models
★ 18kHunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kICon. ICLR 2025 - official implementation for "I-Con: A Unifying Framework for Representation Learning"
★ 136FramePack. Lets make video diffusion practical!
★ 17kAwesome-VQVAE. A collection of resources and papers on Vector Quantized Variational Autoencoder (VQ-VAE) and its application
★ 333Recurrent-Parameter-Generation. The official implementation of Recurrent Diffusion for Large-Scale Parameter Generation.
★ 812d-gaussian-splatting. [SIGGRAPH'24] 2D Gaussian Splatting for Geometrically Accurate Radiance Fields
★ 3.3kGenius. [ACL 2025] A Generalizable and Purely Unsupervised Self-Training Framework
★ 72STC-GS. High-Dynamic Radar Sequence Prediction for Weather Nowcasting Using Spatiotemporal Coherent Gaussian Representation
★ 52Vidu4D. [TPAMI 2025, NeurIPS 2024] Video4DGen: Enhancing Video and 4D Generation through Mutual Optimization
★ 371STeP. STeP: a general and scalable framework for solving video inverse problems with spatiotemporal diffusion priors
★ 32FAR. Code for: "Long-Context Autoregressive Video Modeling with Next-Frame Prediction"
★ 311NBP. Official implementation of Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
★ 42in-context-learning. Jupyter Notebook
★ 245Neural-Network-Diffusion. We introduce a novel approach for parameter generation, named neural network parameter diffusion (p-diff), which employs a standard latent diffusion model to synthesize a new set of parameters
★ 886VEnhancer. Official codes of VEnhancer: Generative Space-Time Enhancement for Video Generation
★ 577Enhance-A-Video. Enhance-A-Video: Better Generated Video for Free
★ 602Dynamic3DGaussians. Python
★ 2.3kc3dgs. Compressed 3D Gaussian Splatting for Accelerated Novel View Synthesis
★ 400Compact-3DGS. The official repository of Compact 3D Gaussian Representation for Radiance Field
★ 5003DGS-Benchmarks.
★ 33gaussian-splatting. Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
★ 23knerf. Code release for NeRF (Neural Radiance Fields)
★ 11kAwesome-Unified-Multimodal. 📖 This is a repository for organizing papers, codes, and other resources related to unified multimodal models.
★ 365DreamScene360. [ECCV 2024] DreamScene360: Unconstrained Text-to-3D Scene Generation with Panoramic Gaussian Splatting
★ 130feature-3dgs. [CVPR 2024 Highlight] Feature 3DGS: Supercharging 3D Gaussian Splatting to Enable Distilled Feature Fields
★ 676Qwen2.5-Omni. Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.
★ 4.1kUnified-MoE-Compression. The official implementation of the paper "Towards Efficient Mixture of Experts: A Holistic Study of Compression Techniques (TMLR)".
★ 89CogVideo. text and image to video generation: CogVideoX (2024) and CogVideo (ICLR 2023)
★ 13kawesome-visual-tokenizer. [WIP🚧] 2025 up-to-date list of resources on visual tokenizers (primarily for visual generation). Give it a star 🌟 if you find it useful.
★ 20Awesome-LLM. Awesome-LLM: a curated list of Large Language Model
★ 27kDreamBooth.github.io. Webpage for DreamBooth
★ 40monst3r. Official Implementation of paper "MonST3R: A Simple Approach for Estimating Geometry in the Presence of Motion"
★ 1.4knerfies.github.io. JavaScript
★ 4.3kIC-Light. More relighting!
★ 8.5kdify. Build Agentic workflows, RAG pipelines, with rich AI model and tool support on one collaborative workspace. Deploy on cloud, VPC, or self-hosted, so teams move from prototype to production without rebuilding the stack.
★ 151kFineParser_CVPR2024. Python
★ 27Awesome-Human-Centric-Foundation-Models. Repo for "Human-Centric Foundation Models: Perception, Generation and Agentic Modeling" (https://arxiv.org/abs/2502.08556)
★ 58HumanOmni. HumanOmni
★ 240Awesome-LLMs-for-Video-Understanding. 🔥🔥🔥 [IEEE TCSVT] Latest Papers, Codes and Datasets on Vid-LLMs.
★ 3.3kSa2VA. Official Repo For Pixel-LLM Codebase: Sa2VA (T-PAMI-26), SAMTok (CVPR-26), VRT (Arxiv-25), SaSaSa2VA (1-st solution for LSVOS)
★ 1.7kR-MeeTo. Give us minutes, we give back a faster Mamba. The official implementation of "Faster Vision Mamba is Rebuilt in Minutes via Merged Token Re-training".
★ 39SEED-Voken. SEED-Voken: A Series of Powerful Visual Tokenizers
★ 1kAutoregressive-Models-in-Vision-Survey. [TMLR 2025🔥] A survey for the autoregressive models in vision.
★ 805LLMSurvey. The official GitHub page for the survey paper "A Survey of Large Language Models".
★ 12kCosmos-Tokenizer. A suite of image and video neural tokenizers
★ 1.7klatex_paper_writing_tips. Tips for Writing a Research Paper using LaTeX
★ 3.8kMeToken. Python
★ 7awesome-MIT-ai-for-climate-change. 🌍 A curated list of MIT faculty that tackle climate change with machine learning for applying students, undergraduates, or others
★ 61Embodied_AI_Paper_List. [Embodied-AI-Survey-2025] Paper List and Resource Repository for Embodied AI
★ 2.1kQuadMamba. Official code for [NeurIPS 2024] QuadMamba: Learning Quadtree-based Selective Scan for Visual State Space Model
★ 46Awesome-Prompt-Adapter-Learning-for-VLMs-CLIP. A curated list of awesome prompt/adapter learning methods for vision-language models like CLIP.
★ 786Awesome-LLM-Strawberry. A collection of LLM papers, blogs, and projects, with a focus on OpenAI o1 🍓 and reasoning techniques.
★ 6.9kAwesome-LWMs. A Collection of Awesome Large Weather Models (LWMs) | AI for Earth (AI4Earth) | AI for Science (AI4Science)
★ 371Awesome-Scientific-Language-Models. A Comprehensive Survey of Scientific Large Language Models and Their Applications in Scientific Discovery (EMNLP'24)
★ 661hydra. Official implementation of "Hydra: Bidirectional State Space Models Through Generalized Matrix Mixers"
★ 175alphafold3-pytorch. Implementation of Alphafold 3 from Google Deepmind in Pytorch
★ 1.7kx-transformers. A concise but complete full-attention transformer with a set of promising experimental features from various papers
★ 5.9kAwesome-Transformer-Attention. An ultimately comprehensive paper list of Vision Transformer/Attention, including papers, codes, and related websites
★ 5.1kAwesome-Visual-Transformer. Collect some papers about transformer with vision. Awesome Transformer with Computer Vision (CV)
★ 3.6kawesome-3D-gaussian-splatting. Curated list of papers and resources focused on 3D Gaussian Splatting, intended to keep pace with the anticipated surge of research in the coming months.
★ 8.8kLearnTrajDep. code for learning trajectory dependencies for human motion prediction
★ 274MambaVision. [CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone
★ 2.2kOMG-Seg. Official Repo For OMG-LLaVA and OMG-Seg codebase [CVPR-24 and NeurIPS-24]
★ 1.4kMG-LLaVA. Official repository for paper MG-LLaVA: Towards Multi-Granularity Visual Instruction Tuning(https://arxiv.org/abs/2406.17770).
★ 160ttt-lm-pytorch. Official PyTorch implementation of Learning to (Learn at Test Time): RNNs with Expressive Hidden States
★ 1.4kLLM-Drop. The official implementation of the paper "Uncovering the Redundancy in Transformers via a Unified Study of Layer Dropping (TMLR)".
★ 191VidToMe. Official Pytorch Implementation for "VidToMe: Video Token Merging for Zero-Shot Video Editing" (CVPR 2024)
★ 2291d-tokenizer. This repo contains the code for 1D tokenizer and generator
★ 1.2kPruner-Zero. [ICML24] Pruner-Zero: Evolving Symbolic Pruning Metric from scratch for LLMs
★ 100MambaTree. [NeurIPS2024 Spotlight] The official implementation of MambaTree: Tree Topology is All You Need in State Space Model
★ 112Vript. Python
★ 161caduceus. Bi-Directional Equivariant Long-Range DNA Sequence Modeling
★ 248awesome-xlstm. A curated list of xLSTM resources
★ 6MLLA. [NeurIPS 2024] Official repository of MLLA
★ 375vision-lstm. xLSTM as Generic Vision Backbone
★ 490flash-linear-attention. 🚀 Efficient implementations for emerging model architectures
★ 5.5kannotated-mamba. Annotated version of the Mamba paper
★ 502Latte. [TMLR 2025] Latte: Latent Diffusion Transformer for Video Generation.
★ 1.9kAwesome-LLM-Prune. Awesome list for LLM pruning.
★ 297awesome-ssm-ml. Reading list for research topics in state-space models
★ 368Mamba_State_Space_Model_Paper_List. [Mamba-Survey-2024] Paper list for State-Space-Model/Mamba and it's Applications
★ 754awesome-time-series-segmentation-papers. This repository contains a reading list of papers on Time Series Segmentation. This repository is still being continuously improved.
★ 546VMamba. VMamba: Visual State Space Models,code is based on mamba
★ 3.2kVim. [ICML 2024] Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
★ 3.9kVision-Mamba-CIFAR10. Python
★ 24PyTorch-VAE. A Collection of Variational Autoencoders (VAE) in PyTorch.
★ 7.7kAwesome-LLM-Robotics. A comprehensive list of papers using large language/multi-modal models for Robotics/RL, including papers, codes, and related websites
★ 4.4kTSFpaper. This repository contains a reading list of papers on Time Series Forecasting/Prediction (TSF) and Spatio-Temporal Forecasting/Prediction (STF). These papers are mainly categorized according to the type of model.
★ 3.2kDiffusion. Minimal multi-gpu implementation of Diffusion Models with Classifier-Free Guidance (CFG)
★ 65Time-Series-Library. A Library for Advanced Deep Time Series Models for General Time Series Analysis.
★ 13kvit-pytorch. Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
★ 25kFontDiffuser. [AAAI2024] FontDiffuser: One-Shot Font Generation via Denoising Diffusion with Multi-Scale Content Aggregation and Style Contrastive Learning
★ 540SimpleCVPaperReading. :smile:博客论文列表:分系列整理
★ 385