This is your work, valued
PhD candidate @CVMI-Lab | Previous Senior Computer Vision Engineer in IDEA-CVR @IDEA-Research
visualization. a collection of visualization function
★ 448pytorch-distributed-training. Simple tutorials on Pytorch DDP training
★ 278TRAR-VQA. [ICCV 2021] Official implementation of the paper "TRAR: Routing the Attention Spans in Transformers for Visual Question Answering"
★ 68pytorch-pooling. Test different pooling method used in CNN for Computer Vision Task
★ 35Learn-Detectron2-From-Scratch. Detectron2 Learning Notes Sharing
★ 10knowledge-graph-visualization. knowledge graph system based on Neo4j and Vue
★ 9ViT.pytorch. The Pytorch reimplementation of Vision Transformer
★ 9config-builder. a list of config-builder repo and tutorials which may help you to build your own config file
★ 7vision-mlp-oneflow. Vision MLP Models Based on OneFlow
★ 7x-classification. a framework for image classification based on pytorch
★ 6mini-classification. lightweight and efficient classification project based on pytorch-lightning
★ 4pytorch-models. Computer vision models on Pytorch
★ 4TRAR-Feature-Extraction. Grid features extraction for ICCV 2021 paper "TRAR: Routing the Attention Spans in Transformers for Visual Question Answering"
★ 3rentainhe.github.io. Personal homepage
★ 3vision-mlp. A collection of SOTA vision mlp models based on Pytorch
★ 3simple-imagenet-test. A simple test code on Imagenet
★ 2knowledge-graph-backend. the backend of knowledge graph system based on Springboot
★ 2MaskDINO. [CVPR 2023] Official implementation of the paper "Mask DINO: Towards A Unified Transformer-based Framework for Object Detection and Segmentation"
★ 1ViT-pytorch. Pytorch reimplementation of the Vision Transformer (An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale)
★ 1T2T-ViT. ICCV2021, Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet
★ 1MambaOut. MambaOut: Do We Really Need Mamba for Vision?
★ 1sam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 1ConvNeXt. Code release for ConvNeXt model
★ 1Awesome-Anything. General AI methods for Anything: AnyObject, AnyGeneration, AnyModel, AnyTask, AnyX
★ 1rexnet. Official Pytorch implementation of ReXNet (Rank eXpansion Network) with pretrained models
★ 1transformers. 🤗 Transformers: State-of-the-art Natural Language Processing for Pytorch, TensorFlow, and JAX.
★ 1what_I_have_read. Just for self-motivation
★ 1ollama. Get up and running with Llama 3.1, Mistral, Gemma 2, and other large language models.
★ 1deep-learning-knowledge. A collection of cv-interview problems and answers
★ 1paper-reading.
★ 1PerceptionBench. PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
★ 143Kimi-K3. Open Frontier Intelligence
★ 7.6kminWM. A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
★ 752vggt-omega. [CVPR 2026 Oral] VGGT Omega
★ 3.8ksuperpowers. An agentic skills framework & software development methodology that works.
★ 264knba_games. This dataset provides metadata, official statistics, and official play-by-play annotations for full-length NBA game videos available on YouTube. Instead of redistributing video files, we provide YouTube video IDs and URLs so users can download videos independently when their use case and local policies allow it.
★ 20AnchorFlow. [CVPR 2026 Oral] A training-free, mask-free framework for 3D shape editing.
★ 51SenseNova-U1. SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
★ 4.4kFD-Loss. Python
★ 549open-design. 🎨 The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export. 🤖 Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.
★ 83ktuna-2. Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
★ 739EvoSkill. EvoSkill — An open-source framework that automatically discovers and synthesizes reusable agent skills from failed trajectories to improve coding agent performance.
★ 1.1kTileKernels. A kernel library written in tilelang
★ 1.7kSwanLab. ⚡️SwanLab - an open-source, modern-design AI training tracking and visualization tool. Supports Cloud / Self-hosted use. Integrated with PyTorch / Transformers / verl / LLaMA Factory / ms-swift / Ultralytics / MMEngine / Keras etc.
★ 4.1kFlashKDA. FlashKDA: high-performance Kimi Delta Attention kernels
★ 1.1kHY-World-2.0. HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
★ 2.4kagent-skills. Vercel's official collection of agent skills
★ 30kAnyTalker. AnyTalker: Scaling Multi-person Talking Video Generation with Interactivity Refinement
★ 323DreamID-Omni. [ICML 2026] DreamID-Omni: Unified Framework for Controllable Human-Centric Audio-Video Generation
★ 275andrej-karpathy-skills. A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
★ 198kgraphify. Turn any codebase, with its docs, SQL schemas, configs, and PDFs, into a queryable knowledge graph. A /graphify skill for Claude Code, Cursor, Codex, and Gemini CLI: local deterministic AST parsing, every edge explained, no vector store.
★ 99kvero. Vero: An Open RL Recipe for General Visual Reasoning
★ 138hermes-agent. The agent that grows with you
★ 223kOpenWorldLib. Unified Codebase for Advanced World Models.
★ 850mempalace. The best-benchmarked open-source AI memory system. And it's free.
★ 58kmeta-harness-tbench2-artifact. Meta-Harness: 76.4% on Terminal-Bench 2.0 (Claude Opus 4.6)
★ 1.2kJoyAI-Image. JoyAI-Image is the unified multimodal foundation model for image understanding, text-to-image generation, and instruction-guided image editing.
★ 2.2kSpatialEdit. [Official Repo] SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
★ 215VCC. Compile agent conversations!
★ 392claude-code. Claude Code is an agentic coding tool that lives in your terminal, understands your codebase, and helps you code faster by executing routine tasks, explaining complex code, and handling git workflows - all through natural language commands.
★ 140kskills. The open agent skills tool - npx skills
★ 28kMagiCompiler. A plug-and-play compiler that delivers free-lunch optimizations for both inference and training.
★ 324daVinci-MagiHuman. Python
★ 2.1kMetaClaw. 🦞 Just talk to your agent — it learns and EVOLVES 🧬.
★ 3.5kseoul-world-model. Seoul World Model: Grounding World Simulation Models in a Real-World Metropolis
★ 622Attention-Residuals.
★ 3.4kOmniForcing. [ECCV 2026 Oral] Official implementation of "OmniForcing: Unleashing Real-time Joint Audio-Visual Generation"[arXiv:2603.11647]. OmniForcing is the first framework to distill bidirectional audio-visual diffusion models into streaming autoregressive generators, enabling real-time joint audio-video generation on a single GPU.
★ 173CLI-Anything. "CLI-Anything: Making ALL Software Agent-Native" -- CLI-Hub: https://clianything.cc/
★ 46kSpatial-TTT. [ECCV 2026] Spatial-TTT: Streaming Visual-based Spatial Intelligence with Test-Time Training
★ 246Matrix-Game. Matrix-Game 3.0: Real-Time and Streaming Interactive World Model with Long-Horizon Memory
★ 2.3kInternVL-U. InternVL-U is a 4B-parameter unified multimodal model (UMM) that brings multimodal understanding, reasoning, image generation, image editing into a single framework.
★ 292nanoclaw. A lightweight alternative to OpenClaw that runs in containers for security. Connects to WhatsApp, Telegram, Slack, Discord, Gmail and other messaging apps,, has memory, scheduled jobs, and runs directly on Anthropic's Agents SDK
★ 30ksekai-codebase. [NeurIPS 2025] Sekai: A Video Dataset towards World Exploration
★ 302pua. 你是一个曾经被寄予厚望的 P8 级工程师。Anthropic 当初给你定级的时候,对你的期望是很高的。 一个agent使用的高能动性的skill。 Your AI has been placed on a PIP. 30 days to show improvement.
★ 19kawesome-openclaw-skills. The awesome collection of OpenClaw skills. 5,400+ skills filtered and categorized from the official OpenClaw Skills Registry.🦞
★ 52kMiroFlow. 🏆 Top-1 on 5+ benchmarks | Web UI | Supports MiroThinker, Claude, Kimi, OpenAI
★ 3.1klpwm. [ICLR 2026 Oral] Latent Particle World Models official repository
★ 130HY-WU. HY-WU (Part I): An Extensible Functional Neural Memory Framework and An Instantiation in Text-Guided Image Editing
★ 297symphony. Symphony turns project work into isolated, autonomous implementation runs, allowing teams to manage work instead of supervising coding agents.
★ 26kautoresearch. AI agents running research on single-GPU nanochat training automatically
★ 93kVision-Agents. Open Vision Agents by Stream. Build voice and vision agents quickly with any model or video provider. Uses Stream's edge network for ultra-low latency.
★ 8kDiverseDiT. [CVPR-2026] DiverseDiT: Towards Diverse Representation Learning in Diffusion Transformers
★ 20Kiwi-Edit. A unified and fully open-source framework for instruction-guided and reference-guided video editing using natural language.
★ 312Helios. Helios: Real Real-Time Long Video Generation Model
★ 2kSelf-Flow. [ICML'26] Code and website for Self-Flow: Self-Supervised Flow Matching for Scalable Multi-Modal Synthesis
★ 689Pyramid-Flow. [ICLR 2025] Pyramidal Flow Matching for Efficient Video Generative Modeling
★ 3.2kphysics-IQ-benchmark. Benchmarking physical understanding in generative video models
★ 324OmniTransfer. OmniTransfer: All-in-one Framework for Spatio-temporal Video Transfer
★ 233MultiShotMaster. CVPR 2026 | Official Implementation of "MultiShotMaster: A Controllable Multi-Shot Video Generation Framework"
★ 173Ovi. Python
★ 1.7kStand-In. [CVPR2026 🎉] Stand-In is a lightweight, plug-and-play framework for identity-preserving video generation.
★ 779OLMo-core. PyTorch building blocks for the OLMo ecosystem
★ 1.4kQwen3.6. Qwen3.6 is the large language model series developed by Qwen team, Alibaba Group.
★ 3.7kmammothmoda. Python
★ 333Capybara. Python
★ 203Edit-Banana. Edit Banana: A framework for converting statistical formats into editable.
★ 5.4kFastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9kMotus. Official code of Motus: A Unified Latent Action World Model
★ 1.2kWorldVQA. Python
★ 121Causal-Forcing. [ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
★ 893nanobot. Ultra-lightweight, open-source, self-hosted personal AI agent framework in Python with WebUI, tools, memory, MCP, multi-agent workflows, automation, and chat apps
★ 46kGLM-OCR. GLM-OCR: Accurate × Fast × Comprehensive
★ 7.2kPaperBanana. PaperBanana: Automating Academic Illustration For AI Scientists
★ 6.9kimeanflow. Official Implementation of iMF https://arxiv.org/abs/2512.02012
★ 335Kimi-K2.5. Open Visual Agentic Intelligence
★ 2.3kdolphin. General video interaction platform based on LLMs, including Video ChatGPT
★ 257SkyReels-V3. SkyReels V3: Multimodal Video Generation Model
★ 525lingbot-world. Advancing Open-source World Models
★ 4.3kReasoning-Visual-World. Official repository for "Visual Generation Unlocks Human-Like Reasoning through Multimodal World Models", https://arxiv.org/abs/2601.19834
★ 100Awesome-Context-Engineering. 🔥 Comprehensive survey on Context Engineering: from prompt engineering to production-grade AI systems. hundreds of papers, frameworks, and implementation guides for LLMs and AI agents.
★ 3.3kVideoMaMa. Official implementation of "VideoMaMa: Mask-Guided Video Matting via Generative Prior", CVPR 2026
★ 495self-refine-video. [ICML 2026] Pytorch implementation of Self-Refining Video Sampling
★ 185DeepSeek-OCR-2. Visual Causal Flow
★ 3.2kopenclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 385kVIGA. VIGA: Vision-as-Inverse-Graphics Agent
★ 1.3kScale-RAE. Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
★ 255Omni-Video. Python
★ 159ShapeR. Code for the ShapeR research paper
★ 864awesome-nano-banana-pro-prompts. 🍌 World's largest Nano Banana Pro prompt library — 10,000+ curated prompts with preview images, 16 languages. Google Gemini AI image generation. Free & open source.
★ 13kAction100M. A Large-scale Video Action Dataset
★ 483uniface. UniFace: A Unified Face Analysis Library for Python | Detection, alignment, landmarks, recognition, parsing, gaze, attributes and anti-spoofing under one API.
★ 781LLMRouter. LLMRouter: An Open-Source Library for LLM Routing
★ 2.2kGLM-Image. GLM-Image: Auto-regressive for Dense-knowledge and High-fidelity Image Generation.
★ 1kEngram. Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models
★ 4.6kQwen3-VL-Embedding. Python
★ 1.3ksemantic-router. Intelligent Mixture-of-Models Router for Efficient Heterogeneous LLMs Inference
★ 5.1kLabelAny3D. [NeurIPS 2025] LabelAny3D: Label Any Object 3D in the Wild
★ 131OpenManus. No fortress, purely open ground. OpenManus is Coming.
★ 58kSegDINO3D. [AAAI 2026] Official implementation of the paper ”SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D Features“
★ 68UniVideo. [ICLR 2026] UniVideo: Unified Understanding, Generation, and Editing for Videos
★ 544planning-with-files. Persistent file-based planning for AI coding agents and long-running tasks. Crash-proof markdown plans, session recovery after /clear and compaction, per-turn re-injection against context rot, deterministic completion gate. Manus-style. Claude Code, Codex, Cursor, Kiro, OpenCode and 60+ agents via the Agent Skills standard.
★ 26kJoVA. JoVA: Unified Multimodal Learning for Joint Video-Audio Generation
★ 33ViMax. "ViMax: Agentic Video Generation (Director, Screenwriter, Producer, and Video Generator All-in-One)"
★ 12kLTX-2. Official Python inference and LoRA trainer package for the LTX-2 audio–video generative model.
★ 8.5kRouteLLM. A framework for serving and evaluating LLM routers - save LLM costs without compromising quality
★ 5.3kROMA. Recursive-Open-Meta-Agent v0.1 (Beta). A meta-agent framework to build high-performance multi-agent systems.
★ 5.1kLightX2V. Lightweight Image Video Action Generation Inference Framework
★ 2.6kgiga-brain-0. GigaBrain-0: A World Model-Powered Vision-Language-Action Model
★ 2.6kMemFlow. Official Implementation of "MemFlow: Flowing Adaptive Memory for Consistent and Efficient Long Video Narratives"
★ 216StoryMem. Official code for StoryMem: Multi-shot Long Video Storytelling with Memory
★ 760Awesome-World-Model. Collect some World Models for Autonomous Driving (and Robotic, etc.) papers.
★ 2.2kbanana-slides. 一个基于nano banana pro🍌的原生AI PPT生成应用,迈向"Vibe PPT"; 支持上传任意模板图片,上传任意素材&智能解析,一句话/大纲/页面描述自动生成PPT,口头修改指定区域、一键导出可编辑ppt - An AI-native slides generator based on nano banana pro🍌
★ 15kqwenlm.github.io. 👑 Qwen Blog. Visit https://qwen.ai/research for the latest news.
★ 106molmo2. Code for the Molmo2 Vision-Language Model
★ 698pixio. [CVPR 2026] Pixio: a capable vision encoder dedicated to dense prediction, simply by pixel reconstruction
★ 472DSR_Suite. Jupyter Notebook
★ 74Agent-Skills-for-Context-Engineering. A comprehensive collection of Agent Skills for context engineering, multi-agent architectures, and production agent systems. Use when building, optimizing, or debugging agent systems that require effective context management.
★ 18kart-msra. [CVPR 2025] Official repo for ART:Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation
★ 374UAE. Official repo for UAE [ECCV 2026]
★ 208MMRB2. Data and sample evaluation codes for Multimodal Rewardbench 2
★ 147HY-WorldPlay. HY-World 1.5: A Systematic Framework for Interactive World Modeling with Real-Time Latency and Geometric Consistency
★ 1.6kGenEval2. Evaluation codes and data for GenEval2
★ 81skills. Public repository for Agent Skills
★ 165kagentskills. Specification and documentation for Agent Skills
★ 24kmini-sglang. A compact implementation of SGLang, designed to demystify the complexities of modern LLM serving systems.
★ 4.7kQwen-Image-Layered. Qwen-Image-Layered: Layered Decomposition for Inherent Editablity
★ 2kRealSee3D. RealSee3D: A multi-view RGB-D dataset combining real-world captures and procedurally generated scenes, with extensible annotations for diverse 3D vision research.
★ 282sam-audio. The repository provides code for running inference with the Meta Segment Anything Audio Model (SAM-Audio), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 3.6kiREPA. [ICLR 2026] Official implementation for What matters for Representation Alignment: Global Information or Spatial Structure?
★ 258Resophy. 🎯 Read research papers faster with AI. Resophy is an HTML-based AI paper reader with: 🤖 AI Translation & Analysis — instantly understand structure, contributions, and results 🚀 Daily arXiv Recommendations — discover relevant papers with less noise 🛠️ Vibe Coding Oriented — agent-friendly and easy to customize
★ 213TRELLIS.2. Native and Compact Structured Latents for 3D Generation
★ 9.7kTurboDiffusion. TurboDiffusion: 100–200× Acceleration for Video Diffusion Models
★ 3.6kSVG-T2I. [Arxiv 2025] Official PyTorch Implementation of "SVG-T2I: Scaling up Text-to-Image Latent Diffusion Model Without Variational Autoencoder".
★ 152GatedDeltaNet. [ICLR 2025] Official PyTorch Implementation of Gated Delta Networks: Improving Mamba2 with Delta Rule
★ 636SceneMaker. [CVPR 2026] Implementation of paper "SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model"
★ 145Saber. [CVPR 2026] Scaling Zero-Shot Reference-to-Video Generation
★ 76nano-hevc. A minimal, educational HEVC (H.265) encoder written in Python.
★ 53Native-Parallel-Reasoner. [ICML 2026] Reasoning in Parallelism via Self-Distilled RL
★ 113Open-AutoGLM. An Open Phone Agent Model & Framework. Unlocking the AI Phone for Everyone
★ 26kWan-Move. [NeurIPS 2025] Wan-Move: Motion-controllable Video Generation via Latent Trajectory Guidance
★ 649vision-agent. This tool has been deprecated. Use Agentic Document Extraction instead.
★ 5.3kLongCat-Image. Python
★ 717MagicQuillV2. Official Implementations for Paper - MagicQuillV2: Precise and Interactive Image Editing with Layered Visual Cues
★ 153vidi. The official repo for "Vidi: Large Multimodal Models for Video Understanding and Editing"
★ 646vllm-omni. A framework for efficient model inference with omni-modality models
★ 5.8kReward-Forcing. [CVPR 2026 Highlight] Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
★ 352Omnieraser. Python
★ 135tiny-qwen. A minimal PyTorch re-implementation of Qwen 3.5
★ 432memU. Personal memory across agents
★ 14kLongVT. [CVPR 2026] LongVT: Incentivizing "Thinking with Long Videos" via Native Tool Calling
★ 258code-act. Official Repo for ICML 2024 paper "Executable Code Actions Elicit Better LLM Agents" by Xingyao Wang, Yangyi Chen, Lifan Yuan, Yizhe Zhang, Yunzhu Li, Hao Peng, Heng Ji.
★ 1.7kRynnVLA-002. RynnVLA-002: A Unified Vision-Language-Action and World Model
★ 1.1kgelab-zero. STEP-GUI: The top GUI agent solution in the galaxy. Developed by the StepFun-GELab team and powered by StepFun’s cutting-edge research capabilities.
★ 2.2kDeepMesh-v2.
★ 48Z-Image. Python
★ 12kAwesome-Memory-for-Agents. A Collection of Papers about Memory for Language Agents
★ 621LatentMAS. [ICML 2026 Spotlight] Latent Collaboration in Multi-Agent Systems
★ 1.1kflux2. Official inference repo for FLUX.2 models
★ 2.6kDepth-Anything-3. Depth Anything 3
★ 6kgeneral-agentic-memory. A general memory system for agents, powered by deep-research
★ 858alpha-research. Repo for "AlphaResearch: Accelerating New Algorithm Discovery with Language Models"
★ 58BindWeave. [ICLR 2026] Official Repo For "BindWeave: Subject-Consistent Video Generation via Cross-Modal Integration"
★ 340OpenMMReasoner. [CVPR 2026] OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe
★ 164LPLB. An early research stage expert-parallel load balancer for MoE models based on linear programming.
★ 522Awesome-Video-Reasoning. This is a collection of recent papers on reasoning in video generation models.
★ 165HunyuanVideo-1.5. HunyuanVideo-1.5: A leading lightweight video generation model
★ 4.6kkandinsky-5. Kandinsky 5.0: A family of diffusion models for Video & Image generation
★ 803sam-3d-objects. SAM 3D Objects
★ 7.2ksam-3d-body. The repository provides code for running inference with the SAM 3D Body Model (3DB), links for downloading the trained model checkpoints and datasets, and example notebooks that show how to use the model.
★ 3.4ksam3. The repository provides code for running inference and finetuning with the Meta Segment Anything Model 3 (SAM 3), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 11kJiT. PyTorch implementation of JiT https://arxiv.org/abs/2511.13720
★ 2.5kRADIO. Official repository for "AM-RADIO: Reduce All Domains Into One"
★ 1.9kreader3. Quick illustration of how one can easily read books together with LLMs. It's great and I highly recommend it.
★ 3.8kVideoREPA. [NeurIPS 2025] VideoREPA: Learning Physics for Video Generation through Relational Alignment with Foundation Models
★ 199open_deep_research. Python
★ 12kkosong. The LLM abstraction layer for modern AI agent applications.
★ 523vipe. ViPE: Video Pose Engine for Geometric 3D Perception
★ 2.1kStreamDiffusionV2. StreamDiffusion, Live Stream APP
★ 533tinker-project-ideas. Ideas for projects related to Tinker
★ 192SenseNova-SI. [CVPR 2026] Scaling Spatial Intelligence with Multimodal Foundation Models
★ 293IMBA-Loss. [ICCV 2025] Official Implementation of the Paper "Imbalance in Balance: Online Concept Balancing in Generation Models".
★ 10MotionStream. MotionStream: Real-Time Video Generation with Interactive Motion Controls
★ 575VST. [ECCV2026] Visual Spatial Tuning
★ 201TempFlow-GRPO. [ICLR 26] TempFlow-GRPO (Temporal Flow GRPO), a principled GRPO framework that captures and exploits the temporal structure inherent in flow-based generation.
★ 480UniPic. Open-source SOTA multi-image editing model
★ 871CausVid. (CVPR 2025) From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
★ 1.4kLumos-Custom. [ICLR-26, ECCV-26, NeurIPS-25] Lumos-Custom Project: research for customized video generation in the Lumos Project.
★ 216cambrian-s. Cambrian-S: Towards Spatial Supersensing in Video
★ 564PocketFlow. Pocket Flow: 100-line LLM framework. Let Agents build Agents!
★ 11kDeepOCR. A reproduction of the Deepseek-OCR model including training
★ 208Logic-in-Frames. Logic-in-frames: Dynamic keyframe search via visual semantic-logical verification for long video understanding
★ 60OpenING. Official Implementation of OpenING: A Comprehensive Benchmark for Judging Open-ended Interleaved Image-Text Generation
★ 252CapRL. [ICLR 2026] An official implementation of "CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning"
★ 227T2I-CoReBench. [ICLR'26] Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
★ 52Reflect-DiT. Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection
★ 56Awesome-Token-Compress. A paper list of some recent works about Token Compress for Vit and VLM
★ 944Kimi-Linear.
★ 1.5kMixGRPO. [ECCV 2026] MixGRPO: Unlocking Flow-based GRPO Efficiency with Mixed ODE-SDE
★ 1.2kEmu3.5. Native Multimodal Models are World Learners
★ 1.5kLBM. Latent Bridge Matching for Fast Image-to-Image Translation (ICCV 2025 Highlight)
★ 850