This is your work, valued
Research Scientist at Adobe
OWOD. (CVPR 2021 Oral) Open World Object Detection
★ 1.1kAwesome-Novel-Class-Discovery. A list of papers that studies Novel Class Discovery
★ 524Awesome-Layout-Generators. An awesome list of layout generation papers
★ 273iOD. (TPAMI 2021) iOD: Incremental Object Detection via Meta-Learning
★ 137ELI. (CVPR 2022) Energy-based Latent Aligner for Incremental Learning
★ 50SDD-Utils. A couple of utilities that can be used with Stanford Drone Dataset
★ 39merlin. (NeurIPS 2020) Meta-Consolidation for Continual Learning
★ 37PyTorch-MAML-and-Reptile. MAML and Reptile sine wave regression example in PyTorch
★ 18PythonBayesianPrediction. Here we are going to evaluate the posterior using MCMC and use it to make Bayesian Prediction.
★ 9DistillGAN. (ICML-W, 2018) Text to image synthesis, by distilling concepts from multiple captions.
★ 3MASON. (ECCV-W 2018) Segmentations using MASON
★ 3Awesome-Mixup. A repository to host recent papers on Manifold Mixup.
★ 3Continual-Learning-101. A repo illustrating the basics of Continual Learning.
★ 3DeepLab. DeepLab Fork from Liang-Chieh Chen's BitBucket DeepLab
★ 2AnnotationEnhancer. A mechanism to automatically remove the false positive annotations and make the bounding boxes even tighter on Stanford Drone Dataset
★ 2UNO. Official implementation of "A Unified Objective for Novel Class Discovery", ICCV2021 (Oral)
★ 2Awesome_Prompting_Papers_in_Computer_Vision. A curated list of prompt-based paper in computer vision and vision-language learning.
★ 1sefa. [CVPR 2021] Closed-Form Factorization of Latent Semantics in GANs
★ 1awesome-diffusion-categorized. collection of diffusion model papers categorized by their subareas
★ 1multimodal-prompt-learning. Official repository of paper titled "MaPLe: Multi-modal Prompt Learning".
★ 1SRCNN-keras. Python
★ 1progressive-gan-pytorch. Implemetatin of Progressive Growing of GANs in PyTorch
★ 1SRCNN-Tensorflow. Image Super-Resolution Using Deep Convolutional Networks in Tensorflow https://arxiv.org/abs/1501.00092v3
★ 1Awesome-Visual-Transformer. Collect some papers about transformer with vision. Awesome Transformer with Computer Vision (CV)
★ 1object-proposals. Repository containing wrapper to obtain various object proposals easily
★ 1Deformable-ConvNets. Deformable Convolutional Networks
★ 1Awesome-Incremental-Learning. Awesome Incremental Learning
★ 1class-incremental-learning. PyTorch implementation of AANets (CVPR 2021) and Mnemonics Training (CVPR 2020)
★ 1LatentMAS. [ICML 2026 Spotlight] Latent Collaboration in Multi-Agent Systems
★ 1.1kAwesome-AI-Memory. TMLR | This survey presents a comprehensive and structured synthesis of memory in LLMs and MLLMs, organizing the literature into a cohesive taxonomy comprising implicit, explicit, and agentic memory paradigms.
★ 37Awesome-Unified-Multimodal. 📖 This is a repository for organizing papers, codes, and other resources related to unified multimodal models.
★ 365star-vector. StarVector is a foundation model for SVG generation that transforms vectorization into a code generation task. Using a vision-language modeling architecture, StarVector processes both visual and textual inputs to produce high-quality SVG code with remarkable precision.
★ 4.5kawesome-weight-space-learning. Awesome papers on weight-space learning
★ 38Liquid. (Accepted by IJCV) Liquid: Language Models are Scalable and Unified Multi-modal Generators
★ 642Kimi-VL. Kimi-VL: Mixture-of-Experts Vision-Language Model for Multimodal Reasoning, Long-Context Understanding, and Strong Agent Capabilities
★ 1.2kfractalgen. PyTorch implementation of FractalGen https://arxiv.org/abs/2502.17437
★ 1.2kOmniParser. A simple screen parsing tool towards pure vision based GUI agent
★ 25kGraphic-design-evaluation. This is the official repository for "Can GPTs Evaluate Graphic Design Based on Design Principles?".
★ 13LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kAutoPresent. Code for the paper "AutoPresent: Designing Structured Visuals From Scratch" (CVPR 2025)
★ 175PosterLLaVA. Python
★ 152vllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kxDiT. xDiT: A Scalable Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
★ 2.7kHunyuanVideo. HunyuanVideo: A Systematic Framework For Large Video Generation Model
★ 12kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kPoetry-In-Pixels-Coling2025. This repository hosts the code and datasets associated with our paper, "Poetry in Pixels: Prompt Tuning for Poem Image Generation via Diffusion Models."
★ 3OmniGen. OmniGen: Unified Image Generation. https://arxiv.org/pdf/2409.11340
★ 4.3kPowerPaint. [ECCV 2024] PowerPaint, a versatile image inpainting model that supports text-guided object inpainting, object removal, image outpainting and shape-guided object inpainting with only a single model. 一个高质量多功能的图像修补模型,可以同时支持插入物体、移除物体、图像扩展、形状可控的物体生成,只需要一个模型
★ 1.1kmochi. The best OSS video generation models, created by Genmo
★ 3.7kLayout-Corrector. Official implementation of "Layout-Corrector: Alleviating Layout Sticking Phenomenon in Discrete Diffusion Model" (ECCV2024)
★ 9POET-continual-action-recognition. Code for the paper: "POET: Prompt Offset Tuning for Continual Human Action Adaptation" (ECCV 2024, Oral)
★ 14MagicQuill. [CVPR'25] Official Implementations for Paper - MagicQuill: An Intelligent Interactive Image Editing System
★ 3.7kQwen3-Coder. Qwen3-Coder is the code version of Qwen3, the large language model series developed by Qwen team.
★ 17kawesome-explainable-cv.
★ 12swarm. Educational framework exploring ergonomic, lightweight multi-agent orchestration. Managed by OpenAI Solution team.
★ 22kautogen. A programming framework for agentic AI
★ 60kCD_CNF. Python
★ 9unsloth. Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek, GLM and other models.
★ 69kMMStar. [NeurIPS 2024] This repo contains evaluation code for the paper "Are We on the Right Way for Evaluating Large Vision-Language Models"
★ 215Awesome-Multimodal-Large-Language-Models. :sparkles::sparkles:Latest Advances on Multimodal Large Language Models
★ 18kunibench. Python Library to evaluate VLM models' robustness across diverse benchmarks
★ 227InternLM. Official release of InternLM series (InternLM, InternLM2, InternLM2.5, InternLM3).
★ 7.3ksam2. The repository provides code for running inference with the Meta Segment Anything Model 2 (SAM 2), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 20kSEED-Story. SEED-Story: Multimodal Long Story Generation with Large Language Model
★ 884smarter. next gen smart vlm reasoner (AI4MATH@ICML'24; Multimodal Algorithmic Reasoning@NeurIPS'24)
★ 7Ranni. Python
★ 237POET_continual_action_recognition. Continual Few-Shot Learning of New Actions With Prompt Tuning
★ 1MiniGPT-5. Official implementation of paper "MiniGPT-5: Interleaved Vision-and-Language Generation via Generative Vokens"
★ 867fiftyone. Refine high-quality datasets and visual AI models
★ 11kZero-Painter. 🔥 [CVPR 2024] The official repo for Zero-Painter!
★ 70InstanceDiffusion. [CVPR 2024] Code release for "InstanceDiffusion: Instance-level Control for Image Generation"
★ 614OpenCOLE. OpenCOLE: Towards Reproducible Automatic Graphic Design Generation [Inoue+, CVPRW2024 (GDUG)]
★ 86LLM-in-Vision. Recent LLM-based CV and related works. Welcome to comment/contribute!
★ 871evals. Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
★ 19kLLM_Categorical_Hierarchical_Representations. Code for 'The Geometry of Categorical and Hierarchical Concepts in Large Language Models' (ICLR 2025, Oral)
★ 115MGM. Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
★ 3.3kRALF. [CVPR 2024 Oral] Official repository for RALF: Retrieval-Augmented Layout Transformer for Content-Aware Layout Generation
★ 143platonic-rep. Python
★ 715axolotl. Go ahead and axolotl questions
★ 12kStoryDiffusion. Accepted as [NeurIPS 2024] Spotlight Presentation Paper
★ 6.4kreka-vibe-eval. Multimodal language model benchmark, featuring challenging examples
★ 189LLaVA-pp. 🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)
★ 841graphist. Official Repo of Graphist
★ 131LibContinual. A Framework of Continual Learning
★ 136amuse. [CVPR 2024] AMUSE: Emotional Speech-driven 3D Body Animation via Disentangled Latent Diffusion
★ 143Control-Color. Control Color: Multimodal Diffusion-based Interactive Image Colorization
★ 197ml-ferret. Python
★ 8.7kGriffon. Official repo of Griffon series including v1(ECCV 2024), v2(ICCV 2025), G, and R, and also the RL tool Vision-R1(CVPR 2026).
★ 250Open-Sora. Open-Sora: Democratizing Efficient Video Production for All
★ 29kHPT. HPT - Open Multimodal LLMs from HyperGAI
★ 313clarity-upscaler. Clarity AI | AI Image Upscaler & Enhancer - free and open-source Magnific Alternative
★ 5.1kGiT. [ECCV2024 Oral🔥] Official Implementation of "GiT: Towards Generalist Vision Transformer through Universal Language Interface"
★ 364LLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kEmu. Emu Series: Generative Multimodal Models from BAAI
★ 1.8kUDiffText. [ECCV 2024] Official repo for UDiffText: A Unified Framework for High-quality Text Synthesis in Arbitrary Images via Character-aware Diffusion Models
★ 236Recommendations-Diffusion-Text-Image. A paper collection of recent diffusion models for text-image generation tasks, e,g., visual text generation, font generation, text removal, text image super resolution, text editing, handwritten generation, scene text recognition and scene text detection.
★ 273PALO. (WACV 2025 - Oral) Vision-language conversation in 10 languages including English, Chinese, French, Spanish, Russian, Japanese, Arabic, Hindi, Bengali and Urdu.
★ 85Awesome-CV-Foundational-Models.
★ 8gemma_pytorch. The official PyTorch implementation of Google's Gemma models
★ 5.7kStableCascade. Official Code for Stable Cascade
★ 6.5kITI-GEN. [ICCV 2023 Oral, Best Paper Finalist] ITI-GEN: Inclusive Text-to-Image Generation
★ 68SUR-adapter. ACM MM'23 (oral), SUR-adapter for pre-trained diffusion models can acquire the powerful semantic understanding and reasoning capabilities from large language models to build a high-quality textual semantic representation for text-to-image generation.
★ 120Energy-Based-CrossAttention. The official repository of "Energy-Based Cross Attention for Bayesian Context Update in Text-to-Image Diffusion Models".
★ 51awesome-diffusion-categorized. collection of diffusion model papers categorized by their subareas
★ 2.2kAwesome-Incremental-Few-Shot-Object-Detection. A paper list for incremental few-shot object detection.
★ 33Awesome-Aesthetics-Assessment. Collection of Aesthetics Assessment Papers for Graphic Designs.
★ 44S3PresignedUpload. Python
★ 1MAG-Edit. MAG-Edit: Localized Image Editing in Complex Scenarios via Mask-Based Attention-Adjusted Guidance (ACM MM2024)
★ 146DS-Fusion. Code for project DS-Fusion
★ 166melfusion. Python
★ 58Image-Color-Aesthetics-and-Quality-Assessment. 🔥[ICCV 2023, Official Code] for paper "Thinking Image Color Aesthetics Assessment: Models, Datasets and Benchmarks". Official Weights and Demos provided. 首个面向图像色彩主观美学评估的数据集、算法和benchmark.
★ 214MagicBrush. [NeurIPS'23] "MagicBrush: A Manually Annotated Dataset for Instruction-Guided Image Editing".
★ 411MiniGPT-4. Open-sourced codes for MiniGPT-4 and MiniGPT-v2 (https://minigpt-4.github.io, https://minigpt-v2.github.io/)
★ 26kWoodpecker. ✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models
★ 649prompt-lookup-decoding. Simple speculative decoding technique, integrated in vLLM and transformers
★ 611Video-LLaVA. PG-Video-LLaVA: Pixel Grounding in Large Multimodal Video Models
★ 264groundingLMM. [CVPR 2024 🔥] Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks.
★ 964pix2video. Code for the paper "Pix2Video: Video Editing using Image Diffusion"
★ 77ControlNet. Let us control diffusion models!
★ 34kLLaVA-Interactive-Demo. LLaVA-Interactive-Demo
★ 380ControlLoRA. ControlLoRA: A Lightweight Neural Network To Control Stable Diffusion Spatial Information
★ 620EditAnything. Edit anything in images powered by segment-anything, ControlNet, StableDiffusion, etc. (ACM MM)
★ 3.4kGrounded-Segment-Anything. Grounded SAM: Marrying Grounding DINO with Segment Anything & Stable Diffusion & Recognize Anything - Automatically Detect , Segment and Generate Anything
★ 18kdiffusers. 🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch.
★ 34ksefa. [CVPR 2021] Closed-Form Factorization of Latent Semantics in GANs
★ 965Asyrp_official. official repo for Asyrp : Diffusion Models already have a Semantic Latent Space (ICLR2023)
★ 289llmblueprint. [ICLR 2024] Official code for the paper "LLM Blueprint: Enabling Text-to-Image Generation with Complex and Detailed Prompts"
★ 85Semantic-SAM. [ECCV 2024] Official implementation of the paper "Semantic-SAM: Segment and Recognize Anything at Any Granularity"
★ 2.9kSoM. [arXiv 2023] Set-of-Mark Prompting for GPT-4V and LMMs
★ 1.6kCVQ-VAE. [ICCV 2023] Online Clustered Codebook
★ 189HiPer. code for HiPer
★ 32DisDiff. [NeurIPS 2023] code for "DisDiff: Unsupervised Disentanglement of Diffusion Probabilistic Models
★ 77CVD-GAN. Python
★ 6DragDiffusion. [CVPR2024, Highlight] Official code for DragDiffusion
★ 1.3kRerender_A_Video. [SIGGRAPH Asia 2023] Rerender A Video: Zero-Shot Text-Guided Video-to-Video Translation
★ 3kAwesome-Generative-Image-Composition. A curated list of papers, code, and resources pertaining to generative image composition or object insertion.
★ 153latent-slot-diffusion. Official Release of NeurIPS 2023 Spotlight paper "Object-Centric Slot Diffusion"
★ 72IDiff-Face. Official repository of the paper: IDiff-Face: Synthetic-based Face Recognition through Fizzy Identity-conditioned Diffusion Models (ICCV 2023)
★ 96dfcil-hgr. [ICCV 2023] Data-Free Class-Incremental Hand Gesture Recognition
★ 17Versatile-Diffusion. Versatile Diffusion: Text, Images and Variations All in One Diffusion Model, arXiv 2022 / ICCV 2023
★ 1.3k3D-OWIS. [NeurIPS2023] 3D-OWIS is capable of detecting unknown instances in inference, and progressively learning novel classes in the process of training.
★ 68piggyback-color. Improved Diffusion-based Image Colorization via Piggybacked Models
★ 72ImageReward. [NeurIPS 2023] ImageReward: Learning and Evaluating Human Preferences for Text-to-image Generation
★ 1.7kFreeU. FreeU: Free Lunch in Diffusion U-Net (CVPR2024 Oral)
★ 1.9kLayoutGPT. Official repo for NeurIPS 2023 paper "LayoutGPT: Compositional Visual Planning and Generation with Large Language Models"
★ 403InST. Official implementation of the paper “Inversion-Based Style Transfer with Diffusion Models” (CVPR 2023)
★ 588LayoutGeneration. Python
★ 209WaveDiff. Official Pytorch Implementation of the paper: Wavelet Diffusion Models are fast and scalable Image Generators (CVPR'23)
★ 442dift. [NeurIPS'23] Emergent Correspondence from Image Diffusion
★ 773Word-As-Image. Python
★ 1.1kAI-Papers-of-the-Week. 🔥Highlighting the top ML papers every week.
★ 13kDLT. Diffusion Layout Transformer implementation.
★ 63Awesome-Open-Vocabulary. (TPAMI 2024) A Survey on Open Vocabulary Learning
★ 999L-CAD. Implementation for "L-CAD: Language-based Colorization with Any-level Descriptions using Diffusion Priors"
★ 45all-seeing. [ICLR 2024 & ECCV 2024] The All-Seeing Projects: Towards Panoptic Visual Recognition&Understanding and General Relation Comprehension of the Open World"
★ 506UnIVAL. [TMLR23] Official implementation of UnIVAL: Unified Model for Image, Video, Audio and Language Tasks.
★ 236riffusion-hobby. Stable diffusion for real-time music generation
★ 3.9kaudiocraft. Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
★ 24kRIVAL. [NeurIPS 2023 Spotlight] Real-World Image Variation by Aligning Diffusion Inversion Chain
★ 154LLMScore. LLMScore: Unveiling the Power of Large Language Models in Text-to-Image Synthesis Evaluation
★ 135Awesome-CV-Foundational-Models.
★ 549muzic. Muzic: Music Understanding and Generation with Artificial Intelligence
★ 4.9kLayoutDETR. The official PyTorch implementation for arXiv'23 paper 'LayoutDETR: Detection Transformer Is a Good Multimodal Layout Designer'
★ 106ODISE. Official PyTorch implementation of ODISE: Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models [CVPR 2023 Highlight]
★ 945ImageBind. ImageBind One Embedding Space to Bind Them All
★ 9.1kdetrex. detrex is a research platform for DETR-based object detection, segmentation, pose estimation and other visual recognition tasks.
★ 2.3krecognize-anything. Open-source and strong foundation image recognition models.
★ 3.7kawesome-segment-anything. Tracking and collecting papers/projects/others related to Segment Anything.
★ 1.7khiera. Hiera: A fast, powerful, and simple hierarchical vision transformer.
★ 1.1kFast-Dawid-Skene. Code for the algorithms in the paper: Vaibhav B Sinha, Sukrut Rao, Vineeth N Balasubramanian. Fast Dawid-Skene: A Fast Vote Aggregation Scheme for Sentiment Classification. KDD WISDOM 2018
★ 45tango. A family of diffusion models for text-to-audio generation.
★ 1.2kDragGAN. Official Code for DragGAN (SIGGRAPH 2023)
★ 36kGroundingDINO. [ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"
★ 10kvisprog. Official code for VisProg (CVPR 2023 Best Paper!)
★ 774AudioLDM. AudioLDM: Generate speech, sound effects, music and beyond, with text.
★ 2.9kInpaint-Anything. Inpaint anything using Segment Anything and inpainting models.
★ 7.7kLaMini-LM. LaMini-LM: A Diverse Herd of Distilled Models from Large-Scale Instructions
★ 822XPretrain. Multi-modality pre-training
★ 511SCUNet. Practical Blind Denoising via Swin-Conv-UNet and Data Synthesis (Machine Intelligence Research 2023)
★ 817MasaCtrl. [ICCV 2023] Consistent Image Synthesis and Editing
★ 843Awesome-Anything. General AI methods for Anything: AnyObject, AnyGeneration, AnyModel, AnyTask, AnyX
★ 1.9kbark. 🔊 Text-Prompted Generative Audio Model
★ 39krich-text-to-image. Rich-Text-to-Image Generation
★ 800msanii. A novel diffusion-based model for synthesizing long-context, high-fidelity music efficiently.
★ 196audio-ai-timeline. A timeline of the latest AI models for audio generation, starting in 2023!
★ 1.9kflex-dm. [CVPR 2023 highlight] Towards Flexible Multi-modal Document Models
★ 59segment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kOpenChatKit. Python
★ 9kCVPR-2023-Papers.
★ 936LayoutDiffusion. diffusion-based layout-to-image generation model
★ 333layout-dm. [CVPR 2023] LayoutDM: Discrete Diffusion Model for Controllable Layout Generation
★ 300MDP-Diffusion. Implementation of MDP: A Generalized Framework for Text-Guided Image Editing by Manipulating the Diffusion Path
★ 67plug-and-play. Official Pytorch Implementation for “Plug-and-Play Diffusion Features for Text-Driven Image-to-Image Translation” (CVPR 2023)
★ 1kpix2pix-zero. Zero-shot Image-to-Image Translation [SIGGRAPH 2023]
★ 1.1kDiffusionDisentanglement. Official implementation of the paper "Uncovering the Disentanglement Capability in Text-to-Image Diffusion Models
★ 174ELITE. ELITE: Encoding Visual Concepts into Textual Embeddings for Customized Text-to-Image Generation (ICCV 2023, Oral)
★ 541CL-DETR. PyTorch implementation of "Continual Detection Transformer for Incremental Object Detection" (CVPR 2023)
★ 112PIDM. Person Image Synthesis via Denoising Diffusion Model (CVPR 2023)
★ 503GracoNet-Object-Placement. [ECCV 2022] Official code for "Learning Object Placement via Dual-path Graph Completion"
★ 102Context-Cluster. [ICLR 2023 Oral] Image as Set of Points
★ 575CLAI_unconf. Website for ContinualAI Unconference
★ 2vision. Datasets, Transforms and Models specific to Computer Vision
★ 18kmachine_unlearning. Existing Literature about Machine Unlearning
★ 967project-webpage. Simple template for a project webpage
★ 34consistency_models. Unofficial Implementation of Consistency Models in pytorch
★ 259reduce_reuse_recycle. ICML 2023: Reduce, Reuse, Recycle: Composing Energy-Based Diffusion Models with MCMC
★ 151blended-latent-diffusion. Official implementation for "Blended Latent Diffusion" [SIGGRAPH 2023]
★ 632mm-cot. Official implementation for "Multimodal Chain-of-Thought Reasoning in Language Models" (stay tuned and more will be updated)
★ 4kk-diffusion. Karras et al. (2022) diffusion models for PyTorch
★ 2.6kmixture-of-diffusers. Mixture of Diffusers for scene composition and high resolution image generation
★ 449MultiDiffusion. Official Pytorch Implementation for "MultiDiffusion: Fusing Diffusion Paths for Controlled Image Generation" presenting "MultiDiffusion" (ICML 2023)
★ 1.1krandom-fourier-features-pytorch. Implementation of "Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains" by Tancik et al.
★ 116direct-inversion. Official code implementation for our paper -- Direct Inversion: Optimization-Free Text-Driven Real Image Editing with Diffusion Models.
★ 27Universal-Guided-Diffusion. Jupyter Notebook
★ 511latex_paper_writing_tips. Tips for Writing a Research Paper using LaTeX
★ 3.8kIllusion-Diffusion. Optical illusions using stable diffusion
★ 225RePaint. Official PyTorch Code and Models of "RePaint: Inpainting using Denoising Diffusion Probabilistic Models", CVPR 2022
★ 2.3kPaint-by-Example. Paint by Example: Exemplar-based Image Editing with Diffusion Models
★ 1.3kEDICT. Jupyter Notebook
★ 322PROB. [CVPR 2023] Official Pytorch code for PROB: Probabilistic Objectness for Open World Object Detection
★ 151nanoGPT. The simplest, fastest repository for training/finetuning medium-sized GPTs.
★ 62kDeepLayout. PyTorch implementation of "LayoutTransformer: Layout Generation and Completion with Self-attention" to appear in ICCV 2021
★ 169character-bert. Main repository for "CharacterBERT: Reconciling ELMo and BERT for Word-Level Open-Vocabulary Representations From Characters"
★ 199diffusion. Denoising Diffusion Probabilistic Models
★ 5.3kAttend-and-Excite. Official Implementation for "Attend-and-Excite: Attention-Based Semantic Guidance for Text-to-Image Diffusion Models" (SIGGRAPH 2023)
★ 770SparK. [ICLR'23 Spotlight🔥] The first successful BERT/MAE-style pretraining on any convolutional network; Pytorch impl. of "Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling"
★ 1.4kRethinking-Text-Segmentation. [CVPR 2021] Rethinking Text Segmentation: A Novel Dataset and A Text-Specific Refinement Approach
★ 275instruct-pix2pix. Python
★ 6.9k