This is your work, valued
Diff-Mosaic. [TGRS 2024]Diff-Mosaic: Augmenting Realistic Representations in Infrared Small Target Detection via Diffusion Prior
★ 20Exploring-Negatives-in-Contrastive-Learning-for-Unpaired-Image-to-Image-Translation. Exploring Negatives in Contrastive Learning for Unpaired Image-to-Image Translation
★ 13Paper-Diffusion-T2M. 个人看的Diffusion model 和 T2M论文的分类
★ 2PhyAgentOS. PhyAgentOS is a self-evolving embodied AI operating system built on agentic workflows.
★ 1.4kPaper-Notes. 📚 数千篇 AI、LLM、NLP、CV 顶会论文解读,每篇 5 分钟读懂核心思想。
★ 1.4kCoRL. Python
★ 9AHAT. Any House Any Task: Scalable Long-Horizon Planning for Abstract Human Tasks
★ 10OPSD. Python
★ 527Awesome-LVLM-Hallucination. up-to-date curated list of state-of-the-art Large vision language models hallucinations research work, papers & resources
★ 325Awesome-Routing-LLMs. A curated list of awesome works in Routing LLMs paradigm (👉 Welcome to submit your contributions to this code repository)
★ 1603DRAE. [Preprint] Any 3D Scene is Worth 1K Tokens: 3D-Grounded Representation for Scene Generation at Scale
★ 57my_codex_skills. This repository collects my personal Codex skills for reusable research workflows.
★ 49Auto-claude-code-research-in-sleep. ARIS ⚔️ (Auto-Research-In-Sleep) — Lightweight Markdown-only skills for autonomous ML research: cross-model review loops, idea discovery, and experiment automation. No framework, no lock-in — works with Claude Code, Codex, OpenClaw, or any LLM agent.
★ 14kEgoThinker. Official implementation of EgoThinker at NIPS 2025
★ 29AAAI26-Exo2Ego. [AAAI 2026] Official Implementation for Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding
★ 3awesome-ego-video-datasets. 🎥 [Awesome] Egocentric / First-Person Video Datasets 📚 Papers, Benchmarks & Resources for Ego Vision
★ 192SkillOrchestra. SkillOrchestra: Learning to Route Agents via Skill Transfer
★ 71OmniXtreme. Python
★ 714awesome-efficient-vla. 🔥 A curated roadmap to the Efficient VLA landscape. We’re keeping this list live—contribute your latest work!
★ 167ReproductionLithoHDwithDL. C++
★ 2E2E-VLA-_AD.
★ 9Video-MME. ✨✨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
★ 788VisualThinker-R1-Zero. Explore the Multimodal “Aha Moment” on 2B Model
★ 624COT_Compresstion_via_Step_entropy. Python
★ 29Awesome-Efficient-R1-style-LRMs.
★ 53Label-Free-RLVR.
★ 311reasoning_loading_bar. Python
★ 56Token_Signature. This repository contains the core implementation of our ICML 2025 paper: "Token Signature: Predicting Chain-of-Thought Gains with Token Decoding Feature in Large Language Models."
★ 44Awesome-RL-based-Reasoning-MLLMs. This repository provides valuable reference for researchers in the field of multimodality, please start your exploratory travel in RL-based Reasoning MLLMs!
★ 1.4kAwesome-Efficient-Reasoning-Models. [TMLR 2025] Efficient Reasoning Models: A Survey
★ 318BC-IB. [ICML 2025] Rethinking Latent Redundancy in Behavior Cloning: An Information Bottleneck Approach for Robot Manipulation
★ 52remix. Python
★ 45Awesome-Efficient-Reasoning. Paper list for Efficient Reasoning.
★ 899BudgetGuidance. [ACL'26 Findings] Steering LLM Thinking with Budget Guidance
★ 33Awesome_Efficient_LRM_Reasoning. 😎 A Survey of Efficient Reasoning for Large Reasoning Models: Language, Multimodality, Agent, and Beyond
★ 357verl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kBrickGPT. [ICCV 2025 Best Paper] Official repository for BrickGPT, the first approach for generating physically stable toy brick models from text prompts.
★ 1.7kthinking-intervention. Used for thinking process intervention of reasoning models such as DeepSeek-R1, effectively controlling the reasoning thinking process. 用于DeepSeek-R1等推理模型的思维过程干预,有效控制推理思考过程
★ 23Dyve. Python
★ 10HaluEval. This is the repository of HaluEval, a large-scale hallucination evaluation benchmark for Large Language Models.
★ 593LLMThinkBench. An Advanced Basic Math Reasoning and Overthinking Evaluation Framework for LLMs
★ 12Awesome-Efficient-Reasoning-LLMs. [TMLR 2025] Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
★ 786EchoWorld. [CVPR 2025] EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
★ 50Awesome-Inference-Time-Scaling. Paper List of Inference/Test Time Scaling/Computing
★ 398compute-optimal-tts. Official codebase for "Can 1B LLM Surpass 405B LLM? Rethinking Compute-Optimal Test-Time Scaling".
★ 288a-m-models. a-m-team's exploration in large language modeling
★ 196Embodied-AI-Guide. [Lumina具身智能社区] 具身智能技术指南 Embodied-AI-Guide
★ 15kRouterEval. A Comprehensive Benchmark for Routing LLMs to Explore Model-level Scaling Up in Large Language Models
★ 121DualPipe. A bidirectional pipeline parallelism algorithm for computation-communication overlap in DeepSeek V3/R1 training.
★ 3kFoundations-of-LLMs. A book for Learning the Foundations of LLMs
★ 17kVTC-CLS. official repo for paper "[CLS] Token Tells Everything Needed for Training-free Efficient MLLMs"
★ 22Awesome-LVLM-paper. :sunglasses: List of papers about Large Multimodal model
★ 30DiffusionNoise. Python
★ 9MirrorDiffusion. zero-shot image-to-image translation, diffusion model, prompt, image-to-image translation, MirrorDiffusion: Stabilizing Diffusion Process in Zero-shot Image Translation by Prompts Redescription and Beyond, IEEE Signal Processing Letters(SPL)
★ 27Diff-Mosaic. [TGRS 2024]Diff-Mosaic: Augmenting Realistic Representations in Infrared Small Target Detection via Diffusion Prior
★ 20MURF. Code of MURF: Mutually Reinforcing Multi-Modal Image Registration and Fusion (IEEE TPAMI 2023)
★ 143MVFA-AD. [CVPR2024 Highlight] Adapting Visual-Language Models for Generalizable Anomaly Detection in Medical Images
★ 248S3Diff. Official implementation of S3Diff
★ 223RFR. Repository for "Infrared Small Target Detection in Satellite Videos: A New Dataset and A Novel Recurrent Feature Refinement Framework"
★ 56MUFusion. (2023' Information Fusion) This is the official implementation for the paper titled "MUFusion: A general unsupervised image fusion network based on memory unit".
★ 38CoCoNet. [IJCV 2024] CoCoNet: Coupled Contrastive Learning Network with Multi-level Feature Ensemble for Multi-modality Image Fusion
★ 119CPR-Coach. Coach-Project
★ 19CPR-CLIP. [IEEE SPL 2023] CPR-CLIP: Multimodal Pre-training for Composite Error Recognition in CPR Training.
★ 8HIT-UAV-Infrared-Thermal-Dataset. A high-altitude infrared thermal dataset for Unmanned Aerial Vehicle-based object detection
★ 260ModelMix. [MICCAI 2024 (early accept)] ModelMix: A New Model-Mixup Strategy to Minimize Vicinal Risk across Tasks for Few-scribble based Cardiac Segmentation
★ 7BasicPBC. Official Implementation of "Learning Inclusion Matching for Animation Paint Bucket Colorization"
★ 306ScaleUpDehazing. Python
★ 9ReVideo. NeurIPS 2024
★ 394CDMamba. Python
★ 115awesome-data-contamination. The Paper List on Data Contamination for Large Language Models Evaluation.
★ 117MemSAM. [CVPR 2024 Oral] MemSAM: Taming Segment Anything Model for Echocardiography Video Segmentation.
★ 195test-time-adaptation. A repository and benchmark for online test-time adaptation.
★ 287Face-Adapter. Python
★ 413rethinking_rotation. [WACV2023] This is the official PyTorch impelementation of our paper "[Rethinking Rotation in Self-Supervised Contrastive Learning: Adaptive Positive or Negative Data Augmentation](https://arxiv.org/abs/2210.12681)"
★ 12img2img-turbo. One-step image-to-image with Stable Diffusion turbo: sketch2image, day2night, and more
★ 2.5kCFAT. [CVPR 2024] CFAT: Unleashing Triangular Windows for Image Super-resolution
★ 93NegativePrompt. The official GitHub page for paper "NegativePrompt: Leveraging Psychology for Large Language Models Enhancement via Negative Emotional Stimuli".
★ 25Stable-Makeup. Pytorch Implementation of "Stable-Makeup: When Real-World Makeup Transfer Meets Diffusion Model" (SIGGRAPH 2025)
★ 229awesome-source-free-test-time-adaptation. Test-time Adaptation, Test-time Training and Source-free Domain Adaptation
★ 549CFDVSR. Collaborative Feedback Discriminative Propagation for Video Super-Resolution
★ 44Official_Remote_Sensing_Mamba. Official code of Remote Sensing Mamba
★ 348VMamba. VMamba: Visual State Space Models,code is based on mamba
★ 3.2kIP-Adapter. The image prompt adapter is designed to enable a pretrained text-to-image diffusion model to generate images with image prompt.
★ 6.6kSIRST-5K. SIRST-5K: Exploring Massive Negatives Synthesis with Self-supervised Learning for Robust Infrared Small Target Detection
★ 41random_quantize. a novel data augmentation method across data modalities
★ 72DiffSketcher. [NIPS 2023] Official implementation for "DiffSketcher: Text Guided Vector Sketch Synthesis through Latent Diffusion Models" https://arxiv.org/abs/2306.14685
★ 310DenseDiffusion. Official Pytorch Implementation of DenseDiffusion (ICCV 2023)
★ 508ICCV2023-paper-code. ICCV2023论文代码汇总
★ 18AnimateDiff. Official implementation of AnimateDiff.
★ 12kDragonDiffusion. ICLR 2024 (Spotlight)
★ 788EGSDE. Official implementation for "EGSDE: Unpaired Image-to-Image Translation via Energy-Guided Stochastic Differential Equations" (NIPS 2022)
★ 237videocomposer. Official repo for VideoComposer: Compositional Video Synthesis with Motion Controllability
★ 958Prompt-Diffusion. Official PyTorch implementation of the paper "In-Context Learning Unlocked for Diffusion Models"
★ 414IPL-Zero-Shot-Generative-Model-Adaptation. [CVPR 2023] Zero-shot Generative Model Adaptation via Image-specific Prompt Learning
★ 86GenerativeDiffusionPrior. Generative Diffusion Prior for Unified Image Restoration and Enhancement (CVPR2023)
★ 315Diffusion-SpaceTime-Attn. Official implementation of the paper "Harnessing the Spatial-Temporal Attention of Diffusion Models for High-Fidelity Text-to-Image Synthesis"
★ 93HCP-Diffusion. A universal Stable-Diffusion toolbox
★ 909Grounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18kFateZero. [ICCV 2023 Oral] "FateZero: Fusing Attentions for Zero-shot Text-based Video Editing"
★ 1.2kpriorMDM. The official implementation of the paper "Human Motion Diffusion as a Generative Prior"
★ 524DSO-noted-in-chinese. 本项目是对直接法视觉里程计Direct Sparse Odometry的详细中文注释
★ 25Hneg_SRC. Official Pytorch implementation of "Exploring Patch-wise Semantic Relation for Contrastive Learning in Image-to-Image Translation Tasks" (CVPR 2022)
★ 67pix2pix-zero. Zero-shot Image-to-Image Translation [SIGGRAPH 2023]
★ 1.1k