This is your work, valued
Working on 3D perception, vision and language
MSMDFusion. [CVPR 2023] MSMDFusion: Fusing LiDAR and Camera at Multiple Scales with Multi-Depth Seeds for 3D Object Detection
★ 207UniToken. [CVPRW 2025] UniToken is an auto-regressive generation model that combines discrete and continuous representations to process visual inputs, making it easy to integrate both visual understanding and image generation tasks seamlessly.
★ 106Lumen. [NeurIPS 2024] Lumen: a Large multimodal model with versatile vision-centric capabilities
★ 25Transformer-backbone. The reproduce of Transformer architecture in paper "Attention is all your need"
★ 18MORE. [ECCV 2022] MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes official implementation
★ 16automatic-matting. The project aims to extract portrait from a picture automatically
★ 5TV-Net. The official code of MM 2021 paper "Two-stage Visual Cues Enhancement Network for Referring Image Segmentation"
★ 3UESTC_OS_experiment. 电子科技大学操作系统进程与资源管理实验代码(python)
★ 3mcu_design. 三天三夜小组mcu设计
★ 3MDU_preprocess. Creating virtual 2D points via multi-depth projection
★ 1judgeval. The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
★ 1kOpenART. Python
★ 4OSWorld-V2. OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
★ 214RepWAM. Code for RepWAM: World Action Modeling with Representation Visual-Action Tokenizers
★ 59VeriLatent. Official code for Adaptive Inference-Time Scaling via Early-Step Latent Verification for Image Editing
★ 3CUA-Gym. Scalable pipeline for synthesizing verifiable RLVR training data for computer-use agents
★ 180Awesome-RL-GUI-Agents. A curated list of awesome RL in GUI Agent papers
★ 52AgentsMeetRL. Awesome List for Agentic RL
★ 1.7kAwesomeOPD. Awesome List for On-Policy Distillation
★ 774gym-anything. Gym-Anything: Turn any Software into an Agent Environment
★ 265andrej-karpathy-skills. A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.
★ 198khermes-agent. The agent that grows with you
★ 223kmolmoweb. Python
★ 581claw-eval. Claw-Eval is an evaluation harness for evaluating LLM as agents. All tasks verified by humans.
★ 738arxiv-monitor. HTML
★ 2GUI-Libra. Official code for paper "GUI-Libra: Training Native GUI Agents to Reason and Act with Action-aware Supervision and Partially Verifiable RL"
★ 66skills. Public repository for Agent Skills
★ 165kOpenRT. Open-source red teaming framework for MLLMs with 42+ attack methods
★ 258ml-unigen. UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
★ 43FreeInpaint. [AAAI 2026] FreeInpaint: Tuning-free Prompt Alignment and Visual Rationality Enhancement in Image Inpainting
★ 18UniREditBench. [ECCV 2026] Offline implementation of UniREditBench: A Unified Reasoning-based Image Editing Benchmark.
★ 58SCOPE. An implementation of SCOPE: Saliency-Coverage Oriented Token Pruning for Efficient Multimodel LLMs (NeurIPS 2025)
★ 31Metis-RISE. Metis-RISE: RL Incentivizes and SFT Enhances Multimodal Reasoning Model Learning
★ 22OpenCUA. [NeurIPS 2025 Spotlight] OpenCUA: Open Foundations for Computer-Use Agents
★ 806GLM-V. GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
★ 2.4kCL4R1T4S. LEAKED SYSTEM PROMPTS FOR CHATGPT, CLAUDE, GEMINI, GROK, PERPLEXITY, CURSOR, LOVABLE, REPLIT, AND MORE! - AI SYSTEMS TRANSPARENCY FOR ALL! 👐
★ 47kumt. A Pytorch Implementation of Unbiased Mean Teacher for Cross-domain Object Detection (CVPR 2021)
★ 105reka-vibe-eval. Multimodal language model benchmark, featuring challenging examples
★ 189Awesome_Think_With_Images. Resources and paper list for "Thinking with Images for LVLMs". This repository accompanies our survey on how LVLMs can leverage visual information for complex reasoning, planning, and generation.
★ 1.5kControlThinker. ControlThinker: Unveiling Latent Semantics for Controllable Image Generation through Visual Reasoning
★ 10Qwen-AD. Official implementation of "Look Before You Decide: Prompting Active Deduction of MLLMs for Assumptive Reasoning"
★ 5OmniGenBench. Python
★ 13describe-anything. [ICCV 2025] Implementation for Describe Anything: Detailed Localized Image and Video Captioning
★ 1.5kr1_reward. ✨✨ [ICLR 2026] R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
★ 292Awesome-Unified-Multimodal-Models. Awesome Unified Multimodal Models
★ 1.3kSimpleAR. Pytorch implementation for the paper titled "SimpleAR: Pushing the Frontier of Autoregressive Visual Generation"
★ 431FlexVAR. Python
★ 130multimodal_rewardbench. Multimodal RewardBench
★ 68Awesome-MCoT. Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
★ 1kAwesome-RL-based-Reasoning-MLLMs. This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!
★ 1.4kDuMo. [AAAI 2025] DuMo: Dual Encoder Modulation Network for Precise Concept Erasure
★ 10MMoP. Multimodal Mixture of Prompt for Vision-Language Models
★ 7RECE. [ECCV 2024] Reliable and Efficient Concept Erasure of Text-to-Image Diffusion Models
★ 93UniToken. [CVPRW 2025] UniToken is an auto-regressive generation model that combines discrete and continuous representations to process visual inputs, making it easy to integrate both visual understanding and image generation tasks seamlessly.
★ 106open-r1. Fully open reproduction of DeepSeek-R1
★ 26kKimi-k1.5.
★ 3.5kOpenTokenizer. Python
★ 21genesis-world. Simulation platform for general-purpose robotics & embodied AI learning.
★ 30kTokenFlow. [CVPR 2025] 🔥 Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".
★ 464TimeMarker. A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability
★ 107Emu3. Next-Token Prediction is All You Need
★ 2.4kOVRE. Python
★ 23ReToMe-VA. [ACM MM 2024] ReToMe-VA: Recursive Token Merging for Video Diffusion-based Unrestricted Adversarial Attack
★ 14AdvQDet. [ACM MM 2024] AdvQDet: Detecting Query-Based Adversarial Attacks with Adversarial Contrastive Prompt Tuning
★ 9EventHallusion. EventHallusion: Diagnosing Event Hallucinations in Video LLMs
★ 34LAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kLumina-mGPT. Official Implementation of "Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining"
★ 647IDForge. [MM24 Oral] Identity-Driven Multimedia Forgery Detection via Reference Assistance
★ 121SEED-Voken. SEED-Voken: A Series of Powerful Visual Tokenizers
★ 1kchameleon. Repository for Meta Chameleon, a mixed-modal early-fusion foundation model from FAIR.
★ 2.1kOmniTokenizer. [NeurIPS 2024]OmniTokenizer: one model and one weight for image-video joint tokenization.
★ 325FeatUp. Official code for "FeatUp: A Model-Agnostic Frameworkfor Features at Any Resolution" ICLR 2024
★ 1.7kDAC. [CVPR 2024] Official implementation of CVPR 2024 paper: "Doubly Abductive Counterfactual Inference for Text-based Image Editing"
★ 26DailyFood-172.
★ 3Lumen. [NeurIPS 2024] Lumen: a Large multimodal model with versatile vision-centric capabilities
★ 26InstaGen. InstaGen: Enhancing Object Detection by Training on Synthetic Dataset, CVPR2024
★ 91LLaVA-MoLE. Python
★ 10MDU_preprocess. Creating virtual 2D points via multi-depth projection
★ 1Griffon. Official repo of Griffon series including v1(ECCV 2024), v2(ICCV 2025), G, and R, and also the RL tool Vision-R1(CVPR 2026).
★ 250InternLM-XComposer. InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
★ 2.9kGCMA. [ACM MM2023] Code Release of GCMA: Generative Cross-Modal Transferable Adversarial Attacks from Images to Videos
★ 12Mapillary2COCO. Transfer Mapillary Vistas Dataset to Coco format
★ 34Qwen-VL. The official repo of Qwen-VL (通义千问-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
★ 6.7kInstructDiffusion. PyTorch implementation of InstructDiffusion, a unifying and generic framework for aligning computer vision tasks with human instructions.
★ 445DriveAGI. [CVPR 2024 Highlight] GenAD: Generalized Predictive Model for Autonomous Driving
★ 802guidance. A guidance language for controlling large language models.
★ 22kNuScenes-QA. [AAAI 2024] NuScenes-QA: A Multi-modal Visual Question Answering Benchmark for Autonomous Driving Scenario.
★ 240SparseFusion. [ICCV 2023] SparseFusion: Fusing Multi-Modal Sparse Representations for Multi-Sensor 3D Object Detection
★ 272BSC-Attack. [AAAI2022] Code Release of Attacking Video Recognition Models with Bullet-Screen Comments
★ 25Self-Universality. Enhancing the Self-Universality for Transferable Targeted Attacks [CVPR 2023 Paper]
★ 36segment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kSurroundOcc. [ICCV 2023] SurroundOcc: Multi-camera 3D Occupancy Prediction for Autonomous Driving
★ 1.1k3D-Dual-Fusion. [Arxiv 2022] This is the official implementation of 3D Dual-Fusion: Dual-Domain Dual-Query Camera-LiDAR Fusion for 3D Object Detection
★ 78SOLOFusion. Time Will Tell: New Outlooks and A Baseline for Temporal Multi-View 3D Object Detection
★ 271GD-MAE. GD-MAE: Generative Decoder for MAE Pre-training on LiDAR Point Clouds (CVPR 2023)
★ 124occupancy-for-nuscenes. 3D occupancy
★ 409ESOL_WSSS.
★ 14ProposalContrast. This repository contains the PyTorch implementation of the ECCV'2022 paper, ProposalContrast: Unsupervised Pre-training for LiDAR-based 3D Object Detection.
★ 58Occupancy-MAE. Official implementation of our TIV'23 paper: Occupancy-MAE: Self-supervised Pre-training Large-scale LiDAR Point Clouds with Masked Occupancy Autoencoders
★ 282BEV-MAE. [AAAI 2024] BEV-MAE: Bird's Eye View Masked Autoencoders for Point Cloud Pre-training in Autonomous Driving Scenarios
★ 86ViGA. "Video Moment Retrieval from Text Queries via Single Frame Annotation" in SIGIR 2022
★ 68AeDet. AeDet: Azimuth-invariant Multi-view 3D Object Detection, CVPR2023
★ 75PENet_ICRA2021. ICRA 2021 "Towards Precise and Efficient Image Guided Depth Completion"
★ 360MSMDFusion. [CVPR 2023] MSMDFusion: Fusing LiDAR and Camera at Multiple Scales with Multi-Depth Seeds for 3D Object Detection
★ 207MORE. [ECCV 2022] MORE: Multi-Order RElation Mining for Dense Captioning in 3D Scenes official implementation
★ 16MVP. Python
★ 302Transformer-backbone. The reproduce of Transformer architecture in paper "Attention is all your need"
★ 18sgmn. Graph-Structured Referring Expressions Reasoning in The Wild, In CVPR 2020, Oral.
★ 117Scan2Cap. [CVPR 2021] Scan2Cap: Context-aware Dense Captioning in RGB-D Scans
★ 107Heuristic_black_box_adversarial_attack_on_video_recognition_models. Python
★ 9TV-Net. The official code of MM 2021 paper "Two-stage Visual Cues Enhancement Network for Referring Image Segmentation"
★ 3Pointnet_Pointnet2_pytorch. PointNet and PointNet++ implemented by pytorch (pure python) and on ModelNet, ShapeNet and S3DIS.
★ 4.9kCDistNet. Official Pytorch implementations of CDistNet: Perceiving Multi-Domain Character Distance for Robust Text Recognition(IJCV)
★ 105LBYLNet. [CVPR2021] Look before you leap: learning landmark features for one-stage visual grounding.
★ 50AugFPN. source code of AugFPN
★ 177Deformable-DETR. Deformable DETR: Deformable Transformers for End-to-End Object Detection.
★ 4kvit-pytorch. Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
★ 25kvision_transformer. Jupyter Notebook
★ 13kLocalization-VizDoom. C
★ 6attention-is-all-you-need-pytorch. A PyTorch implementation of the Transformer model in "Attention is All You Need".
★ 9.8kUESTC_OS_experiment. 电子科技大学操作系统进程与资源管理实验代码(python)
★ 3fastNLP. fastNLP: A Modularized and Extensible NLP Framework. Currently still in incubation.
★ 3.1kmcu_design. 三天三夜小组mcu设计
★ 3