This is your work, valued
Part of the journey is the end
GCNet. GCNet: Non-local Networks Meet Squeeze-Excitation Networks and Beyond
1.2kVFS. Rethinking Self-Supervised Correspondence Learning: A Video Frame-level Similarity Perspective, in ICCV 2021 (Oral)
148IMProv. IMProv: Inpainting-based Multimodal Prompting for Computer Vision Tasks
57GroupViT. GroupViT: Semantic Segmentation Emerges from Text Supervision
25mmdetection. Open MMLab Detection Toolbox with PyTorch 1.0
6Deformable-Convolution-V2-PyTorch. Deformable ConvNets V2 in PyTorch
6comp4411-impressionist. COMP4411 taught by CK
3WifiP2P-Android-App. HKUST UROP 1100
3ODISE. ODISE: Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models
2RMRune. rune vision testing code
2CogVLM2. GPT4V-level open-source multi-modal model based on Llama3-8B
1RMInfantry. RoboMasters Infantry Embedded System
1OFA-fairseq. fairseq from OFA
1Swin-Transformer. This is an official implementation for "Swin Transformer: Hierarchical Vision Transformer using Shifted Windows".
1prismatic-vlms. A flexible and efficient codebase for training visually-conditioned language models (VLMs)
1lintel. A Python module to decode video frames directly, using the FFmpeg C API.
1