This is your work, valued
MSc in CV@MBZUAI
PixWorld. PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
★ 232OneWorld. OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
★ 58VLPTransferAttack. [ECCV2024] Boosting Transferability in Vision-Language Attacks via Diversification along the Intersection Region of Adversarial Trajectory
★ 32Multimodal-RAG-Survey-For-Document. [ACL2026 Main] Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
★ 14CET6-Online. NKU2023年春软件工程大作业--英语六级考试报考系统
★ 3Algorithm. C++
★ 1VJA. [ICML 26 ORAL] When the Prompt Becomes Visual: Vision-Centric Jailbreak Attacks for Large Image Editing Models
★ 28UniWorld-View. Official implementation of UniWorld-View
★ 84diffusion-forcing. code for "Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion"
★ 1.3kMeanFlowNFT. [arXiv 2026] This is the official PyTorch implementation of "MeanFlowNFT: Bringing Forward-Process RL to Average-Velocity Generators".
★ 78DiffusionNFT. [ICLR 2026 Oral] DiffusionNFT: Online Diffusion Reinforcement with Forward Process
★ 994LHTB. Long Horizon Terminal Benchmark with Dense Reward Grading
★ 333ArtiFixer. Python
★ 573OPSD-V. On-Policy Self-Distillation for Post-Training Few-Step Autoregressive Video Generators
★ 456AlayaWorld. Full-stack open-source interactive long-horizon world model.
★ 789RynnWorld-4D. RynnWorld-4D: 4D Embodied World Models for Robotic Manipulation
★ 73Vidu-S1. Vidu S1: A Real-Time Interactive Video Generation Model
★ 225RoboDojo. RoboDojo Official Repo
★ 313PixWorld. PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space
★ 233RDM. Python
★ 79GLD. Official implementation of "Repurposing Geometric Foundation Models for Multi-view Diffusion"
★ 239PhysisForcing. PhysisForcing: Physics Reinforced World Simulator for Robotic Manipulation
★ 108AlignedNorm. [ICML'2026] Official repository of paper titled "AlignedNorm: Prompting Vision–Language Models via Coupled Prompt Field"
★ 8RATs. Implementation of paper "Playful Agentic Robot Learning"
★ 102DrivingGen. [ICLR 2026] DrivingGen: A Comprehensive Benchmark for Generative Video World Models in Autonomous Driving
★ 35tiptop. Official Repo for TiPToP: A Modular Open-Vocabulary Planning System for Robotic Manipulation
★ 134flashdreams. high-performance inference and serving library for interactive autoregressive video and world models
★ 428Echo-Infinity. Official repo for paper "Echo-Infinity: Learnable Evolving Memory for Real-Time Infinite Video Generation"
★ 105JoyAI-Echo. JoyAI-Echo: Pushing the Frontier of Long Audio-Visual Generation
★ 1.8kREST3D. From a single casual image to a visually consistent and physically stable interactive 3D scene.
★ 241Dataset3D. The first "ImageNet" 3D dataset.
★ 91minWM. A Minimal and Elegant Framework & Tutorial for Real-Time Interactive World Models
★ 749AgentDoG. A Diagnostic Guardrail Framework for AI Agent Safety and Security
★ 671SceneGen. [3DV 2026] "SceneGen: Single-Image 3D Scene Generation in One Feedforward Pass"
★ 385PartCrafter. [NeurIPS 2025] PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
★ 2.5karticraft. An Agentic System for Scalable Articulated 3D Asset Generation
★ 1.4kPartFlow. PartFlow: two-stage image-conditioned 3D editing (inference code)
★ 75T2I-L2P. Code for "L2P: Unlocking Latent Potential for Pixel Generation"
★ 180starVLA. StarVLA: A Lego-like Codebase for Vision-Language-Action Model Developing
★ 3.3kPiD. PiD: Fast and High-Resolution Latent Decoding with Pixel Diffusion
★ 997humanize. From Automated Idea Factory to Realization
★ 1.4kAI-Infra-Auto-Driven-SKILLS. Python
★ 703WoG. [ICML 2026] 🏂 World Guidance: World Modeling in Condition Space for Action Generation
★ 162semantic-wm. repository for training action-conditioned latent diffusion world models for robot video generation
★ 74HiDream-O1-Image. Python
★ 1.5kRollingForcing. [ICLR 2026] Official Repo for Rolling Forcing: Autoregressive Long Video Diffusion in Real Time
★ 449Causal-Forcing. [ICML 2026] Official codebase for "Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video Generation" & Causal Forcing++
★ 887AnyFlow. Flow Map OPD for AnyStep Video Diffusion
★ 400Warp-as-History. Warp-as-History: Generalizable Camera-Controlled Video Generation from One Training Video
★ 224FD-Loss. Python
★ 549Flow-OPD. Official Repo of "Flow-OPD: On-Policy Distillation for Flow Matching Models"
★ 266OPSD. Python
★ 519T2PO. 【ICML2026 Spotlight】 T2PO: Uncertainty-Guided Exploration Control for Stable Multi-Turn Agentic Reinforcement Learning
★ 51SenseNova-U1. SenseNova-U series: Native Unified Paradigm with NEO-unify from the First Principles
★ 4.4ktuna-2. Official implementation of Tuna-2: Pixel Embeddings Beat Vision Encoders for Unified Understanding and Generation
★ 739World-R1. [ICML 2026] World-R1: Reinforcing 3D Constraints for Text-to-Video Generation
★ 411hermes-agent. The agent that grows with you
★ 222kHY-World-2.0. HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds
★ 2.4kDiT360. [CVPR 2026] Official implementation of "DiT360: High-Fidelity Panoramic Image Generation via Hybrid Training".
★ 279lyra. Project Lyra: Open Generative 3D World Models
★ 2.2kFast-dLLM. Fork of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"
★ 1ddtree. Python
★ 390Awesome-Feed-Forward-3D. An curated list for feed-forward 3D scene modeling, including research directions, datasets, and applications.
★ 272DMax. DMax: Aggressive Parallel Decoding for dLLMs
★ 128Uni3C. Uni3C: Unifying Precisely 3D-Enhanced Camera and Human Motion Controls for Video Generation [Siggraph Asian 2025]
★ 560ThinkSafe. Python
★ 21lagernvs. Official code for "LagerNVS Latent Geometry for Fully Neural Real-time Novel View Synthesis" (CVPR 2026)
★ 402worldmesh. [ECCV 2026] WorldMesh: Generating Navigable Multi-Room 3D Scenes via Mesh-Conditioned Image Diffusion
★ 145video-subtitle-remover. 基于AI的图片/视频硬字幕去除、文本水印去除,无损分辨率生成去字幕、去水印后的图片/视频文件。无需申请第三方API,本地实现。AI-based tool for removing hard-coded subtitles and text-like watermarks from videos or Pictures.
★ 12kMultimodal-RAG-Survey-For-Document. [ACL2026 Main] Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
★ 15sglang. SGLang is a high-performance serving framework for large language models and multimodal models.
★ 1ETA. [ICLR 2025] PyTorch Implementation of "ETA: Evaluating Then Aligning Safety of Vision Language Models at Inference Time"
★ 34Omni-WorldBench. A comprehensive benchmark specifically designed to evaluate the interactive response capabilities of world models in 4D settings.
★ 106OneWorld. OneWorld: Taming Scene Generation with 3D Unified Representation Autoencoder
★ 58OpenWorldLib. Unified Codebase for Advanced World Models.
★ 848PixelGen. Official repository for “PixelGen: Improving Pixel Diffusion with Perceptual Loss”
★ 275inspatio-world. Python
★ 959SkillJect. SkillJect: Automating Stealthy Skill-Based Prompt Injection for Coding Agents with Trace-Driven Closed-Loop Refinement
★ 75sglang. SGLang is a high-performance serving framework for large language models and multimodal models.
★ 31kSPRVLA. Python
★ 7worldfm. Python
★ 822RobustVLA. Python
★ 21open-lvsm. Python
★ 41TransferAttack. TransferAttack is a pytorch framework to boost the adversarial transferability for image classification.
★ 481GeoThinker. Python
★ 70openclaw. Your own personal AI assistant. Any OS. Any Platform. The lobster way. 🦞
★ 384kdFactory. Easy and Efficient dLLM Fine-Tuning
★ 261Stream-DiffVSR. The official repository of paper "Stream-DiffVSR: Low-Latency Streamable Video Super-Resolution via Auto-Regressive Diffusion"
★ 310FlashVSR. [CVPR 2026] Towards Real-Time Diffusion-Based Streaming Video Super-Resolution — An efficient one-step diffusion framework for streaming VSR with locality-constrained sparse attention and a tiny conditional decoder.
★ 1.7kInvSR. Arbitrary-steps Image Super-resolution via Diffusion Inversion (CVPR 2025)
★ 1.4kSelf-Forcing. Official codebase for "Self Forcing: Bridging Training and Inference in Autoregressive Video Diffusion" (NeurIPS 2025 Spotlight)
★ 3.5krealtime-video. Krea Realtime 14B. An open-source realtime AI video model.
★ 575ml-sharp. Sharp Monocular View Synthesis in Less Than a Second
★ 8.8kODB-dLLM. Implementation of "Orchestrating Dual-Boundaries: An Arithmetic Intensity Inspired Acceleration Framework for Diffusion Language Models"
★ 22d3LLM. [ICML 2026] d3LLM: Ultra-Fast Diffusion LLM 🚀
★ 148Geoshield. Official PyTorch implementation of AAAI 2026 paper Geoshield
★ 13FreeDave. Free Draft-and-Verification: Toward Lossless Parallel Decoding for Diffusion Large Language Models
★ 23DiRL. Python
★ 165JiT. PyTorch implementation of JiT https://arxiv.org/abs/2511.13720
★ 2.5kSkyReels-V2. SkyReels-V2: Infinite-length Film Generative model
★ 7.3kWan2.1-NABLA. Wan: Open and Advanced Large-Scale Video Generative Models
★ 31GEN3C. [CVPR 2025 Highlight] GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
★ 1.4khle. Humanity's Last Exam
★ 1.6kSMDM. Official PyTorch implementation for ICLR2025 paper "Scaling up Masked Diffusion Models on Text"
★ 385LLaDA. Official PyTorch implementation for "Large Language Diffusion Models"
★ 3.9kDiffuLLaMA. [ICLR2025] DiffuGPT and DiffuLLaMA: Scaling Diffusion Language Models via Adaptation from Autoregressive Models
★ 401Open-dLLM. Open diffusion language model for code generation — releasing pretraining, evaluation, inference, and checkpoints.
★ 643MegaDLMs. GPU-optimized framework for training diffusion language models at any scale. The backend of Quokka, Super Data Learners, and OpenMoE 2 training.
★ 343dLLM-RL. [ICLR 2026] Official code for TraceRL: Revolutionizing post-training for Diffusion LLMs, powering the SOTA TraDo series.
★ 511dInfer. dInfer: An Efficient Inference Framework for Diffusion Language Models
★ 475cambrian-s. Cambrian-S: Towards Spatial Supersensing in Video
★ 564dinov3-finetune. Testing adaptation of the DINOv2/3 encoders for vision tasks with Low-Rank Adaptation (LoRA)
★ 503OmniSafeBench-MM. A Unified Benchmark and Toolbox for Multimodal Jailbreak Attack–Defense Evaluation
★ 75WorldScore. Official implementation for WorldScore: A Unified Evaluation Benchmark for World Generation
★ 302AdaCache. Code for our ICCV 2025 paper "Adaptive Caching for Faster Video Generation with Diffusion Transformers"
★ 172Fast-dLLM. Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"
★ 1.1kMagCache. The official code for NeurIPS 2025 "MagCache: Fast Video Generation with Magnitude-Aware Cache"
★ 275DeepSpeed. DeepSpeed is a deep learning optimization library that makes distributed training and inference easy, efficient, and effective.
★ 43kDiffSynth-Studio. Enjoy the magic of Diffusion models!
★ 13kRAE. Official PyTorch Implementation of "Diffusion Transformers with Representation Autoencoders"
★ 2kFlashWorld. Code for "FlashWorld: High-quality 3D Scene Generation within Seconds" (ICLR 2026 Oral)
★ 8334DNeX. 4DNeX: Feed-Forward 4D Generative Modeling Made Easy
★ 840Zero-to-Wan. A minimalistic, hackable code base to finetune Wan video generation model
★ 49MMaDA. MMaDA - Open-Sourced Multimodal Large Diffusion Language Models (dLLMs with block diffusion, mixed-CoT, unified RL)
★ 1.7kLongLive. Long Video Gen Infrastructure
★ 2.5kVideoX-Fun. 📹 A more flexible framework that can generate videos at any resolution and creates videos from images.
★ 2.2kFastVideo. A unified inference and post-training framework for accelerated video generation.
★ 3.9kcolar. [NeurIPS 2025] Think Silently, Think Fast: Dynamic Latent Compression of LLM Reasoning Chains
★ 97verl-tool. A version of verl to support diverse tool use [TMLR 2026]
★ 1kVTool-R1. [ICLR 2026] "VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use"
★ 200WeKnora. Open-source LLM knowledge platform: turn raw documents into a queryable RAG, an autonomous reasoning agent, and a self-maintaining Wiki.
★ 19kOyster. The Oyster series is a set of safety models developed in-house by Alibaba-AAIG, devoted to building a responsible AI ecosystem. | Oyster 系列是 Alibaba-AAIG 自研的安全模型,致力于构建负责任的 AI 生态。
★ 62rStar. Python
★ 1.4kPREMIR. [EMNLP 2025] The official implementation of "Zero-shot Multimodal Document Retrieval via Cross-Modal Question Generation"
★ 15VRAG. Multimodal Retrieval-augmented Generation Framework Built by Tongyi Lab, Alibaba Group.
★ 971R-Zero. [ICLR2026] codes for R-Zero: Self-Evolving Reasoning LLM from Zero Data (https://www.arxiv.org/pdf/2508.05004)
★ 827Thyme. ✨✨ [ICLR 2026] Think Beyond Images
★ 584OmniSpatial. [ICLR 2026] OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
★ 89LVSM. [ICLR 2025 Oral] Official code for "LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias"
★ 550UniUGG. [ICLR 2026] UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
★ 63HPSv3. Official implementation of HPSv3: Towards Wide-Spectrum Human Preference Score (ICCV2025)
★ 331Bagel-Zebra-CoT. https://huggingface.co/datasets/multimodal-reasoning-lab/Zebra-CoT
★ 137anole. [Extended verision ICLR 2025 Blog Track] Anole: An Open, Autoregressive and Native Multimodal Models for Interleaved Image-Text Generation
★ 842TreeVGR. [ICLR'26] Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology
★ 91dots.vlm1. The official repository of the dots.vlm1 instruct models proposed by rednote-hilab.
★ 289DeepSeek-R1.
★ 92kDeepSeek-V3. Python
★ 104kAurora-perception. Python
★ 50metamorph. Code for MetaMorph Multimodal Understanding and Generation via Instruction Tuning
★ 235Director3D. Code for "Director3D: Real-world Camera Trajectory and 3D Scene Generation from Text" (NeurIPS 2024).
★ 381UnifiedReward. Official implementation of UnifiedReward & [NeurIPS 2025] UnifiedReward-Think & UnifiedReward-Flex
★ 796Qwen-Image. Qwen-Image is a powerful image generation foundation model capable of complex text rendering and precise image editing.
★ 8.2kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kAwesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kVLMEvalKit. Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3kLatentCoT-Horizon. 📖 This is a repository for organizing papers, codes, and other resources related to Latent Reasoning.
★ 406Visual-CoT. [Neurips'24 Spotlight] Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
★ 447YUME. The official code of Yume
★ 679reloc3r. [CVPR 2025] Relative camera pose estimation and visual localization with Reloc3r
★ 321Response-Attack. Official implementation of “Response Attack: Exploiting Contextual Priming to Jailbreak Large Language Models” (AAAI 2026).
★ 37ross. [ICLR'25] Reconstructive Visual Instruction Tuning
★ 135Awesome-Unified-Multimodal-Models. Awesome Unified Multimodal Models
★ 1.3kbev-vae. BEV-VAE: A Unified BEV Representation for Generalizable Driving Scene Synthesis
★ 66gill. 🐟 Code and models for the NeurIPS 2023 paper "Generating Images with Multimodal Language Models".
★ 470GoT. Official repository of "GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing"
★ 317MoVieS. [CVPR 2026] Official implementation of "MoVieS: Motion-Aware 4D Dynamic View Synthesis in One Second".
★ 462siren. Welcome to the official repository for Siren, a project aimed at understanding and mitigating harmful behaviors in large language models (LLMs). This repository contains the resources for reproducing the experiments described in our work.
★ 15GuardReasoner. [ICLR Workshop 2025] An official source code for paper "GuardReasoner: Towards Reasoning-based LLM Safeguards".
★ 176Kimi-K2. Kimi K2 is the large language model series developed by Moonshot AI team
★ 11kOvis. A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
★ 1.5kOmniGen2. OmniGen2: Exploration to Advanced Multimodal Generation. https://arxiv.org/abs/2506.18871
★ 4.1kReasonBrain. 【ICML2026】Reasoning to Edit: Hypothetical Instruction-Based Image Editing with Visual Reasoning
★ 27SplatFlow. [CVPR 2025] Official code for the paper "SplatFlow: Multi-View Rectified Flow Model for 3D Gaussian Splatting Synthesis"
★ 143Ovis-U1. An unified model that seamlessly integrates multimodal understanding, text-to-image generation, and image editing within a single powerful framework.
★ 450x-teaming. Python
★ 67Mirage. [CVPR 2026] Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
★ 294panda-guard. Panda Guard is designed for researching jailbreak attacks, defenses, and evaluation algorithms for large language models (LLMs).
★ 69ActorAttack. Python
★ 134RACE. Python
★ 27JailTrickBench. Bag of Tricks: Benchmarking of Jailbreak Attacks on LLMs. Empirical tricks for LLM Jailbreaking. (NeurIPS 2024)
★ 167IB4LLMs. [NeurIPS'24] Protecting Your LLMs with Information Bottleneck
★ 25AutoDefense. AutoDefense: Multi-Agent LLM Defense against Jailbreak Attacks
★ 68Agent-Smith. [ICML 2024] Agent Smith: A Single Image Can Jailbreak One Million Multimodal LLM Agents Exponentially Fast
★ 123FlipAttack. [ICML 2025] An official source code for paper "FlipAttack: Jailbreak LLMs via Flipping".
★ 179JOOD. [CVPR 2025] Official implementation for JOOD "Playing the Fool: Jailbreaking LLMs and Multimodal LLMs with Out-of-Distribution Strategy"
★ 21Awesome-Jailbreak-on-LLMs. Awesome-Jailbreak-on-LLMs is a collection of state-of-the-art, novel, exciting jailbreak methods on LLMs. It contains papers, codes, datasets, evaluations, and analyses.
★ 1.5kVision-Matters. (ArXiv25) Vision Matters: Simple Visual Perturbations Can Boost Multimodal Math Reasoning
★ 60LLM-CBRN-Risks. [ACL 2025 Findings] The official GitHub repo for the paper "Nuclear Deployed: Analyzing Catastrophic Risks in Decision-making of Autonomous LLM Agents"
★ 21Generalization_unified_VLM. Python
★ 24ASVR. Autoregressive Semantic Visual Reconstruction Helps VLMs Understand Better
★ 190vjepa2. PyTorch code and models for VJEPA2 self-supervised learning from video.
★ 4.4kmust3r. MUSt3R: Multi-view Network for Stereo 3D Reconstruction
★ 3343rgs. Python
★ 94AnySplat. [SIGGRAPH Asia 2025 (ACM TOG)] AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
★ 8993dgs-mcmc. [NeurIPS 2024 Spotlight] Implementation of the paper "3D Gaussian Splatting as Markov Chain Monte Carlo"
★ 676sparf. This is the official code release for SPARF: Neural Radiance Fields from Sparse and Noisy Poses [CVPR 2023-Highlight]
★ 300Bagel. Open-source unified multimodal model
★ 6.1kBrickGPT. [ICCV 2025 Best Paper] Official repository for BrickGPT, the first approach for generating physically stable toy brick models from text prompts.
★ 1.7kAwesome-3D-Scene-Generation. A curated list of awesome 3D scene generation papers. (arXiv 2505.05474)
★ 1.1kDeep-Live-Cam. real time face swap and one-click video deepfake with only a single image
★ 95kedm2. EDM2 and Autoguidance -- Official PyTorch implementation
★ 848EVolSplat. official code of CVPR2025 Evolsplat
★ 89AlphaEdit. AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models, ICLR 2025 (Outstanding Paper)
★ 454MaskDiT. Code for Fast Training of Diffusion Models with Masked Transformers
★ 429FlexWorld. Official PyTorch implementation for "FlexWorld: Progressively Expanding 3D Scenes for Flexiable-View Synthesis".
★ 134WorldGen. 🌍 WorldGen - Generate Any 3D Scene in Seconds
★ 2k