I am a final year Ph.D. student in the Computer Vision Department at MBZUAI, working under the supervision of Dr. Salman Khan and Prof. Fahad Khan.
EdgeNeXt. [CADL'22, ECCVW] Official repository of paper titled "EdgeNeXt: Efficiently Amalgamated CNN-Transformer Architecture for Mobile Vision Applications".
417mvits_for_class_agnostic_od. [ECCV'22] Official repository of paper titled "Class-agnostic Object Detection with Multi-modal Transformer".
314ssl_for_fgvc. Self-Supervised Learning for Fine-Grained Image Categorization
26mdef_detr. Python
11SkipVideoFramesUsingOpenCV. This repository contains two different methods to skip frames using OpenCV. It also provides speed comparison of two methods.
10object_detection_in_aerial_images. The repository contains the code for Object Detection in Aerial Images (iSAID dataset) using Faster RCNN and scale-aware data augmentation (SA-AutoAug).
5social_distancing. This repository implements the social distancing violation detection to reduce the spread of COVID-19 and other diseases by ensuring social distancing in public places, shopping malls, hospitals and restaurants.
3ViFi-CLIP. [CVPR 2023] Official repository of paper titled "Fine-tuned CLIP models are efficient video learners".
2darknet. Windows and Linux version of Darknet Yolo v3 & v2 Neural Networks for object detection (Tensor Cores are used)
2mobilevit. Python
2unetr_plus_plus. UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation
2llama-cookbook. Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services
2lmms-eval. Accelerating the development of large multimodal models (LMMs) with one-click evaluation module - lmms-eval.
2Detic. Code release for "Detecting Twenty-thousand Classes using Image-level Supervision".
1cvat_id_switch_solution. The repository contains the code to solve the id switches of tracks labelled using Intel's CVAT tool.
1groundingLMM. Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks [CVPR 2024].
1easy-faster-rcnn.pytorch. An easy implementation of Faster R-CNN (https://arxiv.org/pdf/1506.01497.pdf) in PyTorch.
1LLaVA-pp. 🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)
1AutoGPT. An experimental open-source attempt to make GPT-4 fully autonomous.
1multimodal-prompt-learning. Official repository of paper titled "MaPLe: Multi-modal Prompt Learning".
1Video-ChatGPT. Python
1DCL. Destruction and Construction Learning for Fine-grained Image Recognition
1fairscale. PyTorch extensions for high performance and large scale training.
1DisciplineViolationDetection. Discipline Anomaly Detection using Audio and Video Processing
1LLaVA-pp-HF-Demo. Python
1PyimagesearchComputerVisionCrashCourse. The source code for 17 days crash course by Pyimagesearch
1EvalAI-Starters. How to create a challenge on EvalAI?
1SwiftFormer. SwiftFormer: Efficient Additive Attention for Transformer-based Real-time Mobile Vision Applications
1pytorch-stacked-hourglass. A PyTorch toolkit for 2D Human Pose Estimation.
1detectron2. Detectron2 is FAIR's next-generation platform for object detection, segmentation and other visual recognition tasks.
1object-centric-ovd. Official repository of paper titled "Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection".
1Restormer. [CVPR 2022--Oral] Restormer: Efficient Transformer for High-Resolution Image Restoration. SOTA for motion deblurring, image deraining, denoising (Gaussian/real data), and defocus deblurring.
1