This is your work, valued
Reasearch associate with USTC, Suzhou, School of Biomedical Engineering. Graduated Ph.D. student supervised by Prof. Feng Jiashi in ECE, NUS.
TokenLabeling. Pytorch implementation of "All Tokens Matter: Token Labeling for Training Better Vision Transformers"
★ 436DR-Learning-for-3D-Face. Implementation for paper "Disentangled Representation Learning for 3D Face Shape" CVPR 2019
★ 231Adaptive-Attention. Python
★ 30GrabCut. using opencv-python cv2.grabCut to cut image interactively
★ 10CT-reconstruction. complement of Radom transform in MATLAB
★ 2Mask_Inpainter-. Remove the target object in the picutre
★ 2BookList. 好多书带不走啦,希望有人能继续好好利用它们。
★ 1Fashion-Miner. application of community mining and some algorithm
★ 1Matrix. some useful algorithm about matrix
★ 1RIMD_Reconstruct. C++
★ 1FuseCPath. Official implementation of the paper: Fusion of Multi-scale Heterogeneous Pathology Foundation Models for Whole Slide Image Analysis.
★ 21DeepSeek-OCR. Contexts Optical Compression
★ 24kU-Bench. U-Bench: A Comprehensive Understanding of U-Net through 100-Variant Benchmarking
★ 192HunyuanImage-3.0. HunyuanImage-3.0: A Powerful Native Multimodal Model for Image Generation
★ 3.2kMMR1. [CVPR 2026] MMR1: Enhancing Multimodal Reasoning with Variance-Aware Sampling and Open Resources
★ 217browser-use. 🌐 Make websites accessible for AI agents. Automate tasks online with ease.
★ 107kMedRAX. MedRAX: Medical Reasoning Agent for Chest X-ray - ICML 2025
★ 1.2kLLaMA-Mesh. Unifying 3D Mesh Generation with Language Models
★ 1.2kPuppeteer. [NeurIPS 2025 Spotlight] Official repository for “Puppeteer: Rig and Animate Your 3D Models”
★ 412MedEvalKit. MedEvalKit: A Unified Medical Evaluation Framework
★ 247AgentLaboratory. Agent Laboratory is an end-to-end autonomous research workflow meant to assist you as the human researcher toward implementing your research ideas
★ 5.8kWorldGen. 🌍 WorldGen - Generate Any 3D Scene in Seconds
★ 2kWonderWorld. Code release for https://kovenyu.com/WonderWorld/
★ 7394K4DGen. Python
★ 28vggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kBagel. Open-source unified multimodal model
★ 6.1kAA-CLIP. The official implementation of AA-CLIP: Enhancing Zero-shot Anomaly Detection via Anomaly-Aware CLIP
★ 267CrossMAE. Official Implementation of the CrossMAE paper: Rethinking Patch Dependence for Masked Autoencoders
★ 135cube. Roblox Foundation Model for 3D Intelligence
★ 1.2kGEN3C. [CVPR 2025 Highlight] GEN3C: 3D-Informed World-Consistent Video Generation with Precise Camera Control
★ 1.4kVisual-RFT. Official repository of 'Visual-RFT: Visual Reinforcement Fine-Tuning' & 'Visual-ARFT: Visual Agentic Reinforcement Fine-Tuning'’
★ 2.3kPostoMETRO-Paper. [WACV 2025] PostoMETRO: Pose Token Enhanced Mesh Transformer for Robust 3D Human Mesh Recovery
★ 8LGM. [ECCV 2024 Oral] LGM: Large Multi-View Gaussian Model for High-Resolution 3D Content Creation.
★ 2.1kMVImgNet. CVPR2023 | MVImgNet: A Large-scale Dataset of Multi-view Images
★ 488Vim. [ICML 2024] Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
★ 3.9kR1-Onevision. R1-onevision, a visual language model capable of deep CoT reasoning.
★ 581fractalgen. PyTorch implementation of FractalGen https://arxiv.org/abs/2502.17437
★ 1.2kECAMP. The official implementation of "ECAMP: Entity-centered Context-aware Medical Vision Language Pre-training"
★ 48Inf-Net. Inf-Net: Automatic COVID-19 Lung Infection Segmentation from CT Images, IEEE TMI 2020.
★ 357Opus-MT. Open neural machine translation models and web services
★ 838CopyTranslator. 🔠Foreign language reading and translation assistant based on copy and translate.
★ 18kGPT_API_free. Free ChatGPT&DeepSeek API Key,免费ChatGPT&DeepSeek API。免费接入DeepSeek API和GPT4 API,支持 gpt | deepseek | claude | gemini | grok 等排名靠前的常用大模型。
★ 39kFooocus. Focus on prompting and generating
★ 52kLLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25kBard-API. The unofficial python package that returns response of Google Bard through cookie value.
★ 5.2kMM-Vet. MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities (ICML 2024)
★ 331ms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15kAnimatedDrawings. Code to accompany "A Method for Animating Children's Drawings of the Human Figure"
★ 13kPlatypus. Code for fine-tuning Platypus fam LLMs using LoRA
★ 625neuralangelo. Official implementation of "Neuralangelo: High-Fidelity Neural Surface Reconstruction" (CVPR 2023)
★ 4.6kFastChat. An open platform for training, serving, and evaluating large language models. Release repo for Vicuna and Chatbot Arena.
★ 40kLAVIS. LAVIS - A One-stop Library for Language-Vision Intelligence
★ 11kXrayGPT. [BIONLP@ACL 2024] XrayGPT: Chest Radiographs Summarization using Medical Vision-Language Models.
★ 530Huatuo-Llama-Med-Chinese. Repo for BenCao [original name: HuaTuo (华驼)], Instruction-tuning Large Language Models with Chinese Medical Knowledge. 本草(原名:华驼)模型仓库,基于中文医学知识的大语言模型指令微调
★ 5kopen_flamingo. An open-source framework for training large multimodal models.
★ 4.1kmed-flamingo. Python
★ 452DragDiffusion. [CVPR2024, Highlight] Official code for DragDiffusion
★ 1.3kMetaGPT. 🌟 The Multi-Agent Framework: First AI Software Company, Towards Natural Language Programming
★ 70kEmu. Emu Series: Generative Multimodal Models from BAAI
★ 1.8kChatLaw. ChatLaw:A Powerful LLM Tailored for Chinese Legal. 中文法律大模型
★ 7.6kprolificdreamer. ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score Distillation (NeurIPS 2023 Spotlight)
★ 1.6kPanoHead. Code Repository for CVPR 2023 Paper "PanoHead: Geometry-Aware 3D Full-Head Synthesis in 360 degree"
★ 2kFastSAM. Fast Segment Anything
★ 8.4kLOMO. LOMO: LOw-Memory Optimization
★ 993CoLLiE. Collaborative Training of Large Language Models in an Efficient Way
★ 419ChatPaper. Use ChatGPT to summarize the arXiv papers. 全流程加速科研,利用chatgpt进行论文全文总结+专业翻译+润色+审稿+审稿回复
★ 20kMatting-Anything. Matting Anything Model (MAM), an efficient and versatile framework for estimating the alpha matte of any instance in an image with flexible and interactive visual or linguistic user prompt guidance.
★ 715threestudio. A unified framework for 3D content generation.
★ 7kHuatuoGPT. HuatuoGPT, Towards Taming Language Models To Be a Doctor. (An Open Medical GPT)
★ 1.3kqlora. QLoRA: Efficient Finetuning of Quantized LLMs
★ 11kEVA3D. [ICLR 2023 Spotlight] EVA3D: Compositional 3D Human Generation from 2D Image Collections
★ 601voice-changer. リアルタイムボイスチェンジャー Realtime Voice Changer
★ 21kMedical-SAM-Adapter. A lightweight adapter bridges SAM with medical imaging [MedIA]
★ 1.3kshap-e. Generate 3D objects conditioned on text or images
★ 12kMultimodal-GPT. Multimodal-GPT
★ 1.5kunlimiformer. Public repo for the NeurIPS 2023 paper "Unlimiformer: Long-Range Transformers with Unlimited Length Input"
★ 1.1kAudioGPT. AudioGPT: Understanding and Generating Speech, Music, Sound, and Talking Head
★ 10kzero123. Zero-1-to-3: Zero-shot One Image to 3D Object (ICCV 2023)
★ 3.1kdinov2. PyTorch code and models for the DINOv2 self-supervised learning method.
★ 13kAnything-3D. Segment-Anything + 3D. Let's lift anything to 3D.
★ 1.6kAGIEval. Python
★ 774Open-Assistant. OpenAssistant is a chat-based assistant that understands tasks, can interact with third-party systems, and retrieve information dynamically to do so.
★ 37kAutoGPT. AutoGPT is the vision of accessible AI for everyone, to use and to build on. Our mission is to provide the tools, so that you can focus on what matters.
★ 186kHumanSD. [ICCV 2023] The official implementation of paper "HumanSD: A Native Skeleton-Guided Diffusion Model for Human Image Generation"
★ 305Painter. Painter & SegGPT Series: Vision Foundation Models from BAAI
★ 2.6ksegment-anything. The repository provides code for running inference with the SegmentAnything Model (SAM), links for downloading the trained model checkpoints, and example notebooks that show how to use the model.
★ 55kTune-A-Video. [ICCV 2023] Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
★ 4.4kinceptionnext. InceptionNeXt: When Inception Meets ConvNeXt (CVPR 2024)
★ 353ChatDoctor. Python
★ 3.6kSparK. [ICLR'23 Spotlight🔥] The first successful BERT/MAE-style pretraining on any convolutional network; Pytorch impl. of "Designing BERT for Convolutional Networks: Sparse and Hierarchical Masked Modeling"
★ 1.4kevals. Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.
★ 19kTaskMatrix. Python
★ 34kBioGPT. Python
★ 4.5kFedCLS. [ICLR2023] Towards Understanding and Mitigating Dimensional Collapse in Heterogeneous Federated Learning (https://arxiv.org/abs/2210.00226)
★ 40torchscale. Foundation Architecture for (M)LLMs
★ 3.1kDiffCloth. Code repository for our paper DiffCloth: Differentiable Cloth Simulation with Dry Frictional Contact
★ 431ENARF-GAN. Python
★ 71pyAudioAnalysis. Python Audio Analysis Library: Feature Extraction, Classification, Segmentation and Applications
★ 6.3knice-slam. [CVPR'22] NICE-SLAM: Neural Implicit Scalable Encoding for SLAM
★ 1.6kstable-dreamfusion. Text-to-3D & Image-to-3D & Mesh Exportation with NeRF + Diffusion.
★ 8.8kD-NeRF. Jupyter Notebook
★ 591body-model-visualizer. GUI for visualization and interactive editing of SMPL-family body models ie. SMPL, SMPL-X, MANO, FLAME.
★ 469insetGAN. Official repository of CVPR 2022 paper InsetGAN
★ 169taming-transformers. Taming Transformers for High-Resolution Image Synthesis
★ 6.5kxrnerf. OpenXRLab Neural Radiance Field (NeRF) Toolbox and Benchmark
★ 586yolov7. Implementation of paper - YOLOv7: Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors
★ 14kbdd100k. Toolkit of BDD100K Dataset for Heterogeneous Multitask Learning - CVPR 2020 Oral Paper
★ 563CLRNet. Pytorch implementation of our paper "CLRNet: Cross Layer Refinement Network for Lane Detection" (CVPR2022 Acceptance).
★ 590OpenPCDet. OpenPCDet Toolbox for LiDAR-based 3D Object Detection.
★ 5.7kobjsdf. :t-rex: [ECCV‘22] Pytorch implementation of 'Object-Compositional Neural Implicit Surfaces'
★ 190AvatarGen. [Preprint] AvatarGen: A 3D Generative Model for Animatable Human Avatars
★ 247RegionCLIP. [CVPR 2022] Official code for "RegionCLIP: Region-based Language-Image Pretraining"
★ 816GIRAFFEHD. Official Github repository for the CVPR 2022 paper "GIRAFFE HD: A High-Resolution 3D-aware Generative Model"
★ 70SemanticStyleGAN. release code for SemanticStyleGAN (CVPR 2022)
★ 273GANet. A Keypoint-based Global Association Network for Lane Detection. Accepted by CVPR 2022
★ 262stylegan-xl. [SIGGRAPH'22] StyleGAN-XL: Scaling StyleGAN to Large Diverse Datasets
★ 995HybridNets. HybridNets: End-to-End Perception Network
★ 683TokenLabeling. Pytorch implementation of "All Tokens Matter: Token Labeling for Training Better Vision Transformers"
★ 436StyleGAN-Human. StyleGAN-Human: A Data-Centric Odyssey of Human Generation
★ 1.2kCwD. Official Implementation of CVPR 2022 paper: "Mimicking the Oracle: An Initial Phase Decorrelation Approach for Class Incremental Learning"
★ 35sem2nerf. 😺 [ECCV'22] Sem2NeRF: Converting Single-View Semantic Masks to NeRFs
★ 126TensoRF. [ECCV 2022] Tensorial Radiance Fields, a novel approach to model and reconstruct radiance fields
★ 1.2kexpose. ExPose - EXpressive POse and Shape rEgression
★ 668ICON. [CVPR'22] ICON: Implicit Clothed humans Obtained from Normals
★ 1.7kpytorch-openpose. pytorch implementation of openpose including Hand and Body Pose Estimation.
★ 2.3kPoint-MAE. [ECCV2022] Masked Autoencoders for Point Cloud Self-supervised Learning
★ 638grid_sample1d. pytorch cuda extension of grid_sample1d
★ 49StyleSDF. Python
★ 535DirectVoxGO. Direct voxel grid optimization for fast radiance field reconstruction.
★ 1.1kinstant-ngp. Instant neural graphics primitives: lightning fast NeRF and more
★ 18ktorch-ngp. A pytorch CUDA extension implementation of instant-ngp (sdf and nerf), with a GUI.
★ 2.2keg3d. Python
★ 3.3kstylegan3. Official PyTorch implementation of StyleGAN3
★ 6.9kCIPS-3D. 3D-aware GANs based on NeRF (arXiv).
★ 609open_clip. An open source implementation of CLIP.
★ 14kDIT. [TNNLS] Full Transformer Framework for Robust Point Cloud Registration with Deep Information Interaction
★ 39DietNeRF. Python
★ 138svox2. Plenoxels: Radiance Fields without Neural Networks
★ 2.9kibot. iBOT :robot:: Image BERT Pre-Training with Online Tokenizer (ICLR 2022)
★ 778Real-ESRGAN. Real-ESRGAN aims at developing Practical Algorithms for General Image/Video Restoration.
★ 36kpoolformer. PoolFormer: MetaFormer Is Actually What You Need for Vision (CVPR 2022 Oral)
★ 1.4kppuda. Code for Parameter Prediction for Unseen Deep Architectures (NeurIPS 2021)
★ 491Text-Image-Augmentation. Geometric Augmentation for Text Image
★ 494lama. 🦙 LaMa Image Inpainting, Resolution-robust Large Mask Inpainting with Fourier Convolutions, WACV 2022
★ 10kPedestron. [Pedestron] Generalizable Pedestrian Detection: The Elephant In The Room. @ CVPR2021
★ 699vilbert_beta. Jupyter Notebook
★ 478pnp-detr. Implementation of ICCV21 paper: PnP-DETR: Towards Efficient Visual Analysis with Transformers
★ 120AdelaiDet. AdelaiDet is an open source toolbox for multiple instance-level detection and recognition tasks.
★ 3.5kSynthText. Code for generating synthetic text images as described in "Synthetic Data for Text Localisation in Natural Images", Ankush Gupta, Andrea Vedaldi, Andrew Zisserman, CVPR 2016.
★ 2.1kgeti. Build computer vision models in a fraction of the time and with less data.
★ 1.3knex-code. Code release for NeX: Real-time View Synthesis with Neural Basis Expansion
★ 605putting-nerf-on-a-diet. Putting NeRF on a Diet: Semantically Consistent Few-Shot View Synthesis Implementation
★ 266nerf. Code release for NeRF (Neural Radiance Fields)
★ 11knerf-pytorch. A PyTorch implementation of NeRF (Neural Radiance Fields) that reproduces the results.
★ 6.1kYOLOX. YOLOX is a high-performance anchor-free YOLO, exceeding yolov3~v5 with MegEngine, ONNX, TensorRT, ncnn, and OpenVINO supported. Documentation: https://yolox.readthedocs.io/
★ 11kPretrained-IPT. Python
★ 470ViLT. Code for the ICML 2021 (long talk) paper: "ViLT: Vision-and-Language Transformer Without Convolution or Region Supervision"
★ 1.5kdiffvg. Differentiable Vector Graphics Rasterization
★ 1.3kParlAI. A framework for training and evaluating AI models on a variety of openly available dialogue datasets.
★ 11kalphafold. Open source code for AlphaFold 2.
★ 15kmdetr. Python
★ 1.1kgiraffe. This repository contains the code for the CVPR 2021 paper "GIRAFFE: Representing Scenes as Compositional Generative Neural Feature Fields"
★ 1.2krecovering-unbiased-scene-graphs. Official implementation of "Recovering the Unbiased Scene Graphs from the Biased Ones" (ACMMM 2021)
★ 78pytorch-grad-cam. Advanced AI Explainability for computer vision. Support for CNNs, Vision Transformers, Classification, Object detection, Segmentation, Image similarity and more.
★ 13kAD-NeRF. This repository contains a PyTorch implementation of "AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis".
★ 1.1kAudioCLIP. Source code for models described in the paper "AudioCLIP: Extending CLIP to Image, Text and Audio" (https://arxiv.org/abs/2106.13043)
★ 872lightseq. LightSeq: A High Performance Library for Sequence Processing and Generation
★ 3.3kvolo. VOLO: Vision Outlooker for Visual Recognition
★ 948VisionPermutator. MLP-Like Vision Permutator for Visual Recognition (PyTorch)
★ 192LV-BERT. LV-BERT: Exploiting Layer Variety for BERT (Findings of ACL 2021)
★ 18PoseAug. [CVPR 2021 Best Paper Award Candidate] PoseAug: A Differentiable Pose Augmentation Framework for 3D Human Pose Estimation, (Oral, Best Paper Award Finalist)
★ 383Refiner_ViT. Python
★ 110yolov5. Ultralytics YOLOv5 in PyTorch for object detection, instance segmentation, classification, training, and export.
★ 58kYOLOS. [NeurIPS 2021] You Only Look at One Sequence
★ 902DynamicViT. [NeurIPS 2021] [T-PAMI] DynamicViT: Efficient Vision Transformers with Dynamic Token Sparsification
★ 668Transformer-Explainability. [CVPR 2021] Official PyTorch implementation for Transformer Interpretability Beyond Attention Visualization, a novel method to visualize classifications by Transformer based networks.
★ 2kmmf. A modular framework for vision & language multimodal research from Facebook AI Research (FAIR)
★ 5.6kImageNet21K. Official Pytorch Implementation of: "ImageNet-21K Pretraining for the Masses"(NeurIPS, 2021) paper
★ 779BLIP. Official Implementation of CVPR2021 paper: Continual Learning via Bit-Level Information Preserving
★ 39dino. PyTorch code for Vision Transformers training with the Self-Supervised learning method DINO
★ 7.6kvissl. VISSL is FAIR's library of extensible, modular and scalable components for SOTA Self-Supervised Learning with images.
★ 3.3kSimCSE. [EMNLP 2021] SimCSE: Simple Contrastive Learning of Sentence Embeddings https://arxiv.org/abs/2104.08821
★ 3.7kpytorch-deeplab-xception. DeepLab v3+ model in PyTorch. Support different backbones.
★ 3kPatchVisionTransformer. Jupyter Notebook
★ 74pytorch-image-models. The largest collection of PyTorch image encoders / backbones. Including train, eval, inference, export scripts, and pretrained weights -- ResNet, ResNeXT, EfficientNet, NFNet, Vision Transformer (ViT), MobileNetV4, MobileNet-V3 & V2, RegNet, DPN, CSPNet, Swin Transformer, MaxViT, CoAtNet, ConvNeXt, and more
★ 37krl-generalization-paper. A list of papers regarding generalization in (deep) reinforcement learning
★ 156deepmind-research. This repository contains implementations and illustrative code to accompany DeepMind publications
★ 15kTeacher-free-Knowledge-Distillation. Knowledge Distillation: CVPR2020 Oral, Revisiting Knowledge Distillation via Label Smoothing Regularization
★ 582Hadamard-Matrix-for-hashing. CVPR2020/TNNLS2023: Central Similarity Quantization/Hashing for Efficient Image and Video Retrieval
★ 239AOT. AOT: Appearance Optimal Transport Based Identity Swapping for Forgery Detection (NeurIPS 2020)
★ 39CondenseNetV2. [CVPR 2021] CondenseNet V2: Sparse Feature Reactivation for Deep Networks
★ 86Adaptive-Attention. Python
★ 30dvit_repo. Python
★ 141Swin-Transformer. This is an official implementation for "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows".
★ 16kTensorflow. Jupyter Notebook
★ 14CoordAttention. Code for our CVPR2021 paper coordinate attention
★ 1.2kDALL-E. PyTorch package for the discrete VAE used for DALL·E.
★ 11kvilbert-multi-task. Multi Task Vision and Language
★ 824T2T-ViT. ICCV2021, Tokens-to-Token ViT: Training Vision Transformers from Scratch on ImageNet
★ 1.2krelabel_imagenet. Python
★ 406CLIP. CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
★ 34kyoutube-8m. Starter code for working with the YouTube-8M dataset.
★ 2.4ksam. Python
★ 640vision_transformer. Jupyter Notebook
★ 13kvit-keras. Keras implementation of ViT (Vision Transformer)
★ 352DeFCN. End-to-End Object Detection with Fully Convolutional Network
★ 494flax. Flax is a neural network library for JAX that is designed for flexibility.
★ 7.3kjax. Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more
★ 36kmixreg. Code for our NeurIPS 2020 paper Improving Generalization in Reinforcement Learning with Mixture Regularization
★ 34SimSiam. A pytorch implementation for paper 'Exploring Simple Siamese Representation Learning'
★ 830