This is your work, valued
I am a machine learning researcher at Qualcomm AI Research. I have broad interests in computer vision, computer graphics, and machine learning.
Context-Aware-Matting. The project is the inference implementation of our ICCV2019 paper "Context-aware Image Matting for Simultaneous Foreground and Alpha Estimation"
★ 132VFIPS. The PyTorch implementation for A Perceptual Quality Metric for Video Frame Interpolation
★ 45msspl. The official codes for Fast Monte Carlo Rendering via Multi-Resolution Sampling
★ 15tiny-cnn. deep learning(convolutional neural networks) in C++11/TBB
★ 2nerf. Code release for NeRF (Neural Radiance Fields)
★ 1OpenSpec. Spec-driven development (SDD) for AI coding assistants.
★ 63kSDPO. Reinforcement Learning via Self-Distillation (SDPO)
★ 1kMotus. Official code of Motus: A Unified Latent Action World Model
★ 1.2klingbot-va. [RSS 2026] Causal video-action world model for generalist robot control
★ 1.7kSimpleVLA-RL. [ICLR 2026] SimpleVLA-RL: Scaling VLA Training via Reinforcement Learning
★ 1.8kGDPO. Official implementation of GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
★ 494OpenTinker. OpenTinker is an RL-as-a-Service infrastructure for foundation models
★ 677mup. maximal update parametrization (µP)
★ 1.7kPaCoRe. PaCoRe: Learning to Scale Test-Time Compute with Parallel Coordinated Reasoning
★ 338DeepResearch. Tongyi Deep Research, the Leading Open-source Deep Research Agent
★ 20kDepth-Anything-3. Depth Anything 3
★ 6klejepa. Python
★ 1.3kcosmos-predict2.5. Cosmos-Predict2.5, the latest version of the Cosmos World Foundation Models (WFMs) family, specialized for simulating and predicting the future state of the world in the form of video.
★ 1.3kmap-anything. MapAnything: Universal Feed-Forward Metric 3D Reconstruction
★ 3.6kremake. やり直すんだ。そして、次はうまくやる。
★ 10ksort-free-gs. Re-implementation of Sort-Free Gaussian Splatting via Weighted Sum Rendering
★ 37CUT3R. Official implementation of Continuous 3D Perception Model with Persistent State
★ 1.5knative-sparse-attention-triton. Efficient triton implementation of Native Sparse Attention.
★ 284fast3r. [CVPR 2025] Fast3R: Towards 3D Reconstruction of 1000+ Images in One Forward Pass
★ 1.6kprope. Cameras as Relative Positional Encoding
★ 744AnySplat. [SIGGRAPH Asia 2025 (ACM TOG)] AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
★ 901Direct3D-S2. [NeurIPS 2025] Direct3D‑S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
★ 1.3kLong-LRM. Self-reimplemented version of Long-LRM.
★ 240LVSM. [ICLR 2025 Oral] Official code for "LVSM: A Large View Synthesis Model with Minimal 3D Inductive Bias"
★ 550ttt-video-dit. Official PyTorch implementation of One-Minute Video Generation with Test-Time Training
★ 2.4kvggt. [CVPR 2025 Best Paper Award] VGGT: Visual Geometry Grounded Transformer
★ 14kQwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kGPT-SoVITS. 1 min voice data can also be used to train a good TTS model! (few shot voice cloning)
★ 60kR1-V. Witness the aha moment of VLM with less than $3.
★ 4.1kPIKE-RAG. PIKE-RAG: sPecIalized KnowledgE and Rationale Augmented Generation
★ 2.5kChatDev. ChatDev 2.0: Dev All through LLM-powered Multi-Agent Collaboration
★ 34kLLaVA. [NeurIPS'23 Oral] Visual Instruction Tuning (LLaVA) built towards GPT-4V level capabilities and beyond.
★ 25klmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kTokenFlow. [CVPR 2025] 🔥 Official impl. of "TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation".
★ 464LLM-Reverse-Curriculum-RL. Implementation of the ICML 2024 paper "Training Large Language Models for Reasoning through Reverse Curriculum Reinforcement Learning" presented by Zhiheng Xi et al.
★ 116gsplat-website. Contains website data for https://gsplat.tech
★ 38splat. WebGL 3D Gaussian Splat Viewer
★ 3.1kNoPoSplat. [ICLR'25 Oral] No Pose, No Problem: Surprisingly Simple 3D Gaussian Splats from Sparse Unposed Images
★ 974xformers. Hackable and optimized Transformers building blocks, supporting a composable construction.
★ 11kCompact-3DGS. The official repository of Compact 3D Gaussian Representation for Radiance Field
★ 500gsplat. CUDA accelerated rasterization of gaussian splatting
★ 5.5ktimeful.app. Timeful (formerly Schej) is a scheduling platform helps you find the best time for a group to meet. It is a free availability poll that is easy to use and integrates with your calendar.
★ 1.7kslang-gaussian-rasterization. Slang
★ 376clean-fid. PyTorch - FID calculation with proper image resizing and quantization steps [CVPR 2022]
★ 1.2kdgm-eval. Codebase for evaluation of deep generative models as presented in Exposing flaws of generative model evaluation metrics and their unfair treatment of diffusion models
★ 211DNGaussian. [CVPR'24] DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth Normalization
★ 365multinerf. A Code Release for Mip-NeRF 360, Ref-NeRF, and RawNeRF
★ 3.8ktorch-splatting. A pure pytorch implementation of 3D gaussian Splatting
★ 454BlenderNeRF. Easy NeRF synthetic dataset creation within Blender
★ 1kfwdgrad. Implementation of "Gradients without backpropagation" paper (https://arxiv.org/abs/2202.08587) using functorch
★ 114generative-compression. TensorFlow Implementation of Generative Adversarial Networks for Extreme Learned Image Compression
★ 533SpacetimeGaussians. [CVPR 2024] Spacetime Gaussian Feature Splatting for Real-Time Dynamic View Synthesis
★ 826generative-models. Generative Models by Stability AI
★ 27kDynamic3DGaussians. Python
★ 2.3kslang. Making it easier to work with shaders
★ 5.5kNeuRBF. Python
★ 314gaussian-splatting. Original reference implementation of "3D Gaussian Splatting for Real-Time Radiance Field Rendering"
★ 23kmitsuba3. Mitsuba 3: A Retargetable Forward and Inverse Renderer
★ 2.9kR-FID-Robustness-of-Quality-Measures-for-GANs. This is the official repo for the ECCV paper: "On the Robustness of Quality Measures for GANs
★ 11distill-sd. Segmind Distilled diffusion
★ 618crunch. Advanced DXTc texture compression and transcoding library
★ 900BK-SDM. A Compressed Stable Diffusion for Efficient Text-to-Image Generation [ECCV'24]
★ 321aimet-model-zoo. Python
★ 346RVRT. Recurrent Video Restoration Transformer with Guided Deformable Attention (NeurlPS2022, official repository)
★ 451BasicVSR_PlusPlus. Official repository of "BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment"
★ 812monodepth2. [ICCV 2019] Monocular depth estimation from a single image
★ 4.5kEMA-VFI. [CVPR 2023] Extracting Motion and Appearance via Inter-Frame Attention for Efficient Video Frame Interpolatio
★ 498PCWNet. Python
★ 60PhoenixGo. Go AI program which implements the AlphaGo Zero paper
★ 2.9kAoT_Dataset. CVPR18: Learning and Using the Arrow of Time
★ 40DCVC. Deep Contextual Video Compression
★ 796RealtimeStereo. Attention-Aware Feature Aggregation for Real-time Stereo Matching on Edge Devices (ACCV, 2020)
★ 185MC-Denoising-via-Auxiliary-Feature-Guided-Self-Attention. Official implementation of MC Denoising via Auxiliary Feature Guided Self-Attention (SIGGRAPH Asia 2021 paper)
★ 46vision_blender. A Blender addon for generating synthetic ground truth data for Computer Vision applications
★ 529pygta5. Explorations of Using Python to play Grand Theft Auto 5.
★ 3.9kAwesome-Image-Harmonization. A curated list of papers, code and resources pertaining to image harmonization.
★ 536SimSwap. An arbitrary face-swapping framework on images and videos with one single trained model!
★ 5.2kfaceswap. Deepfakes Software For All
★ 57kstereo-magnification. Code accompanying the SIGGRAPH 2018 paper "Stereo Magnification: Learning View Synthesis using Multiplane Images"
★ 417blender-cli-rendering. Blender Python scripts for rendering images directly from command-line interface
★ 823nvdiffrec. Official code for the CVPR 2022 (oral) paper "Extracting Triangular 3D Models, Materials, and Lighting From Images".
★ 2.3kAirSim. Open source simulator for autonomous vehicles built on Unreal Engine / Unity, from Microsoft AI & Research
★ 18kVSFA. [official] Quality Assessment of In-the-Wild Videos (ACM MM 2019)
★ 217DVQA. Deep learning-based Video Quality Assessment
★ 482tungsten. High performance physically based renderer in C++11
★ 1.8kAI-Paper-Collector. MLNLP社区用来更好进行论文搜索的工具。Fully-automated scripts for collecting AI-related papers
★ 1.2kface-parsing.PyTorch. Using modified BiSeNet for face parsing in PyTorch
★ 2.6kface-seg. Semantic segmentation for hair, face and background
★ 124hair-segmentation. hair segmentation in mobile device
★ 252renderdoc_for_game_data. Access data with annotations from GTA5 using renderdoc
★ 56Video-Frame-Interpolation-Transformer. Python
★ 103HCFlow. Official PyTorch code for Hierarchical Conditional Flow: A Unified Framework for Image Super-Resolution and Image Rescaling (HCFlow, ICCV2021)
★ 195TensoRF. [ECCV 2022] Tensorial Radiance Fields, a novel approach to model and reconstruct radiance fields
★ 1.2kunrealrox-plus. Upgraded version of UnrealROX, as plugin
★ 31gamehook. C++
★ 135FFHQ-Aging-Dataset. FFHQ-Aging Dataset
★ 285ExtraNet. ExtraNet: Real-time Extrapolated Rendering for Low-latency Temporal Supersampling
★ 82revisiting-sepconv. an implementation of Revisiting Adaptive Convolutions for Video Frame Interpolation using PyTorch
★ 93Context-Aware-Matting. The project is the inference implementation of our ICCV2019 paper "Context-aware Image Matting for Simultaneous Foreground and Alpha Estimation"
★ 132instant-ngp. Instant neural graphics primitives: lightning fast NeRF and more
★ 18kraspiquickcv2. Install opencv in minutes on raspbian
★ 4SwinIR. SwinIR: Image Restoration Using Swin Transformer (official repository)
★ 5.6kml-cvnets. CVNets: A library for training computer vision networks
★ 2kkilonerf. Code for KiloNeRF: Speeding up Neural Radiance Fields with Thousands of Tiny MLPs
★ 492hyperstyle. Official Implementation for "HyperStyle: StyleGAN Inversion with HyperNetworks for Real Image Editing" (CVPR 2022) https://arxiv.org/abs/2111.15666
★ 1kSOAT. Official PyTorch repo for StyleGAN of All Trades: Image Manipulation with Only Pretrained StyleGAN.
★ 375SRCNN-pytorch. PyTorch implementation of Image Super-Resolution Using Deep Convolutional Networks (ECCV 2014)
★ 670photometric_optimization. Photometric optimization code for creating the FLAME texture space and other applications
★ 592FastFlowNet. FastFlowNet: A Lightweight Network for Fast Optical Flow Estimation (ICRA 2021)
★ 329VideoPose3D. Efficient 3D human pose estimation in video using 2D keypoint trajectories
★ 4.1kPerceptualSimilarity. LPIPS metric. pip install lpips
★ 4.3kUniPose. We propose UniPose, a unified framework for human pose estimation, based on our “Waterfall” Atrous Spatial Pooling architecture, that achieves state-of-art-results on several pose estimation metrics. Current pose estimation methods utilizing standard CNN architectures heavily rely on statistical postprocessing or predefined anchor poses for joint localization. UniPose incorporates contextual seg- mentation and joint localization to estimate the human pose in a single stage, with high accuracy, without relying on statistical postprocessing methods. The Waterfall module in UniPose leverages the efficiency of progressive filter- ing in the cascade architecture, while maintaining multi- scale fields-of-view comparable to spatial pyramid config- urations. Additionally, our method is extended to UniPose- LSTM for multi-frame processing and achieves state-of-the- art results for temporal pose estimation in Video. Our re- sults on multiple datasets demonstrate that UniPose, with a ResNet backbone and Waterfall module, is a robust and efficient architecture for pose estimation obtaining state-of- the-art results in single person pose detection for both sin- gle images and videos.
★ 223deep-head-pose. :fire::fire: Deep Learning Head Pose Estimation using PyTorch.
★ 1.7karticulated-animation. Code for Motion Representations for Articulated Animation paper
★ 1.3kfirst-order-model. This repository contains the source code for the paper First Order Motion Model for Image Animation
★ 15kDeepV2D. Python
★ 673RAFT-3D. Python
★ 264MichiGAN. MichiGAN: Multi-Input-Conditioned Hair Image Generation for Portrait Editing (SIGGRAPH 2020)
★ 292expansion. Upgrading Optical Flow to 3D Scene Flow through Optical Expansion, CVPR 2020 (Oral).
★ 181Human-Segmentation-PyTorch. Human segmentation models, training/inference code, and trained weights, implemented in PyTorch
★ 574matting_human_datasets. 人像matting数据集,包含34427张图像和对应的matting结果图。
★ 639sbmc. Sample-based Monte Carlo Denoising using a Kernel-Splatting Network [Siggraph 2019]
★ 92BackgroundMattingV2. Real-Time High-Resolution Background Matting
★ 7.2kdeit. Official DeiT repository
★ 4.4kvit-pytorch. Implementation of Vision Transformer, a simple way to achieve SOTA in vision classification with only a single transformer encoder, in Pytorch
★ 25koh-my-userscripts. 我的油猴脚本仓库
★ 4.1knerf. Code release for NeRF (Neural Radiance Fields)
★ 11kkkndme_tianya. 天涯 kkndme 神贴聊房价
★ 19kjonbarron.github.io. HTML
★ 3.6kirr. Iterative Residual Refinement for Joint Optical Flow and Occlusion Estimation (CVPR 2019)
★ 208ScopeFlow. Dynamic Scene Scoping for Optical Flow (CVPR 2020)
★ 96BMBC. BMBC: Bilateral Motion Estimation with Bilateral Cost Volume for Video Interpolation, ECCV 2020
★ 86RAFT. Python
★ 4.1kMaskFlownet-Pytorch. Pytorch implementation of MaskFlownet
★ 88FBA_Matting. Official repository for the paper F, B, Alpha Matting
★ 502MaskFlownet. [CVPR 2020, Oral] MaskFlownet: Asymmetric Feature Matching with Learnable Occlusion Mask
★ 374pytorch-randaugment. Unofficial PyTorch Reimplementation of RandAugment.
★ 635ARFlow. The official PyTorch implementation of the paper "Learning by Analogy: Reliable Supervision from Transformations for Unsupervised Optical Flow Estimation".
★ 269EDSR-PyTorch. PyTorch version of the paper 'Enhanced Deep Residual Networks for Single Image Super-Resolution' (CVPRW 2017)
★ 2.6kSelFlow. SelFlow: Self-Supervised Learning of Optical Flow
★ 400closed-form-matting. Python implementation of A. Levin D. Lischinski and Y. Weiss. A Closed Form Solution to Natural Image Matting. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), June 2006, New York
★ 444softmax-splatting. an implementation of softmax splatting for differentiable forward warping using PyTorch
★ 516DRN. Closed-loop Matters: Dual Regression Networks for Single Image Super-Resolution
★ 440CAIN. Source code for AAAI 2020 paper "Channel Attention Is All You Need for Video Frame Interpolation"
★ 351pytorch_nms. CUDA implementation of NMS for PyTorch
★ 88PySceneDetect. :movie_camera: Python and OpenCV-based scene cut/transition detection program & library.
★ 5.1kmogicians-manual. Flutter version【膜法指南】open source project
★ 518GCNet. GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
★ 1.2kCCNet. CCNet: Criss-Cross Attention for Semantic Segmentation (TPAMI 2020 & ICCV 2019).
★ 1.5kevil-huawei. Evil Huawei - 华为作过的恶
★ 9.2kblender-compile. (NOT MAINTIANED) docker environment for compiling blender 2.8
★ 12DALI. A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.
★ 5.7ktensorflow-densenet. Tensorflow-DenseNet with ImageNet Pretrained Models
★ 167OpenVINO-Custom-Layers. Tutorial for Using Custom Layers with OpenVINO (Intel Deep Learning Toolkit)
★ 106EDVR. Winning Solution in NTIRE19 Challenges on Video Restoration and Enhancement (CVPR19 Workshops) - Video Restoration with Enhanced Deformable Convolutional Networks. EDVR has been merged into BasicSR and this repo is a mirror of BasicSR.
★ 1.6kknn-matting. Python implementation of KNN Matting, CVPR 2012 / TPAMI 2013 http://dingzeyu.li/projects/knn/
★ 125tpu. Reference models and tools for Cloud TPUs.
★ 5.3kdarts. Differentiable architecture search for convolutional and recurrent networks
★ 4kawesome-architecture-search. A curated list of awesome architecture search resources
★ 1.2kpytorch-mobilenet-v3. MobileNetV3 in pytorch and ImageNet pretrained models
★ 806HRNet-Semantic-Segmentation. The OCR approach is rephrased as Segmentation Transformer: https://arxiv.org/abs/1909.11065. This is an official implementation of semantic segmentation for HRNet. https://arxiv.org/abs/1908.07919
★ 3.3kranking. Learning to Rank in TensorFlow
★ 2.8ktensorboardX. tensorboard for pytorch (and chainer, mxnet, numpy, ...)
★ 8kGrid-Anchor-based-Image-Cropping. Project page of the CVPR2019 paper "Reliable and Efficient Image Cropping: A Grid Anchor based Approach"
★ 127ViewProposalNet. VPN for CVPR18
★ 58ViewEvaluationNet. VEN for cvpr2018
★ 37flickr-cropping-dataset. A dataset that includes photos downloaded from Flickr and annotations that indicates a local window representing a good composition.
★ 79AffinityBasedMattingToolbox. A collection of common affinity-based image matting and matte refinement algorithms.
★ 176view-finding-network. A deep ranking network that learns to find good compositions in a photograph.
★ 76PyTorch-YOLOv3. Minimal PyTorch implementation of YOLOv3
★ 7.4kAlignedReID-Re-Production-Pytorch. Reproduce AlignedReID: Surpassing Human-Level Performance in Person Re-Identification, using Pytorch.
★ 645selfconsistency. Code for the paper: Fighting Fake News: Image Splice Detection via Learned Self-Consistency
★ 197DCGAN-tensorflow-slim. Implementation of DCGAN in TensorFlow-Slim
★ 8PWC-Net. PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume, CVPR 2018 (Oral)
★ 1.7ksemantic-segmentation-pytorch. Pytorch implementation for Semantic Segmentation/Scene Parsing on MIT ADE20K dataset
★ 5.1kmodels. Models and examples built with TensorFlow
★ 78kPyTorch-Encoding. A CV toolkit for my papers.
★ 2kDetectron. FAIR's research platform for object detection research, implementing popular algorithms like Mask R-CNN and RetinaNet.
★ 26kpytorch-semseg. Semantic Segmentation Architectures Implemented in PyTorch
★ 3.4ktuio-mouse-driver. TUIO Mouse Driver
★ 9bottom-up-attention. Bottom-up attention model for image captioning and VQA, based on Faster R-CNN and Visual Genome
★ 1.5kDeformable-ConvNets. Deformable Convolutional Networks + MST + Soft-NMS
★ 233improved_wgan_training. Code for reproducing experiments in "Improved Training of Wasserstein GANs"
★ 2.4kXJTU-Thesis-Template. This is a LaTeX template for doctoral thesis of Xi`an Jiaotong University (XJTU). The platform is Windows + TexLive + XeLaTex.
★ 36WassersteinGAN. Python
★ 3.2kfast-pixel-cnn. Speed up PixelCNN++ image generation by up to a 183 times
★ 474Faster-RCNN_TF. Faster-RCNN in Tensorflow
★ 2.3kkeras. Deep Learning for humans
★ 64kdeeplearning-papernotes. Summaries and notes on Deep Learning research papers
★ 4.4kAdversarialNetsPapers. Awesome paper list with code about generative adversarial nets
★ 6.6kWassersteinGAN.tensorflow. Tensorflow implementation of Wasserstein GAN - arxiv: https://arxiv.org/abs/1701.07875
★ 412tensorflow-zh. 谷歌全新开源人工智能系统TensorFlow官方文档中文版
★ 12kiGAN. Interactive Image Generation via Generative Adversarial Networks
★ 4kvae_tutorial. Caffe code to accompany my Tutorial on Variational Autoencoders
★ 523face-py-faster-rcnn. Face Detection with the Faster R-CNN
★ 381caffe. Caffe: a fast open framework for deep learning.
★ 35kcaffe. A fork of Caffe with OpenMPI-based Multi-GPU (mainly data parallel) support for action recognition and more. More documentation please see the original readme.
★ 550XX-Net. A proxy tool to bypass GFW.
★ 33kYCM-Generator. Generates config files for YouCompleteMe (https://github.com/Valloric/YouCompleteMe)
★ 908spf13-vim. The ultimate vim distribution
★ 15k