VisTR. [CVPR2021 Oral] End-to-End Video Instance Segmentation with Transformers
757SOLOv2. SOLOv2: Dynamic, Faster and Stronger, achives 39.5mAP on coco test-dev (36 epochs result)
247PAR. [CVPR2025 Highlight] PAR: Parallelized Autoregressive Visual Generation. https://yuqingwang1029.github.io/PAR-project
186SOLO. SOLO: Segmenting Objects by Locations
164TokenBridge. [ICCV2025] TokenBridge: Bridging Continuous and Discrete Tokens for Autoregressive Visual Generation. https://yuqingwang1029.github.io/TokenBridge
158CondInst. Conditional Convolutions for Instance Segmentation, achives 37.1mAP on coco val
146Multiple-instance-learning. Pytorch implementation of three Multiple Instance Learning or Multi-classification papers
136Yolact_fcos. YOLACT: Real-time Instance Segmentation on the FCOS detector (without bbox cropping), achives 35.2mAP on coco val
91CubiD. [CVPR2026 Highlight] Cubic Discrete Diffusion: Discrete Visual Generation on High-Dimensional Representation Tokens https://arxiv.org/abs/2603.19232
63RepresentationForcing. JavaScript
3Loong-video. [Project page] Generating Minute-level Long Videos with Autoregressive Language Models https://yuqingwang1029.github.io/Loong-video/
1