Abu Dhabi, UAE, San Francisco, USA

Muhammad Maaz

Expert
@mmaaz60

I am a final year Ph.D. student in the Computer Vision Department at MBZUAI, working under the supervision of Dr. Salman Khan and Prof. Fahad Khan.

EdgeNeXt. [CADL'22, ECCVW] Official repository of paper titled "EdgeNeXt: Efficiently Amalgamated CNN-Transformer Architecture for Mobile Vision Applications".

417

mvits_for_class_agnostic_od. [ECCV'22] Official repository of paper titled "Class-agnostic Object Detection with Multi-modal Transformer".

314

ssl_for_fgvc. Self-Supervised Learning for Fine-Grained Image Categorization

26

mdef_detr. Python

11

SkipVideoFramesUsingOpenCV. This repository contains two different methods to skip frames using OpenCV. It also provides speed comparison of two methods.

10

object_detection_in_aerial_images. The repository contains the code for Object Detection in Aerial Images (iSAID dataset) using Faster RCNN and scale-aware data augmentation (SA-AutoAug).

5

social_distancing. This repository implements the social distancing violation detection to reduce the spread of COVID-19 and other diseases by ensuring social distancing in public places, shopping malls, hospitals and restaurants.

3

ViFi-CLIP. [CVPR 2023] Official repository of paper titled "Fine-tuned CLIP models are efficient video learners".

2

darknet. Windows and Linux version of Darknet Yolo v3 & v2 Neural Networks for object detection (Tensor Cores are used)

2

mobilevit. Python

2

unetr_plus_plus. UNETR++: Delving into Efficient and Accurate 3D Medical Image Segmentation

2

llama-cookbook. Welcome to the Llama Cookbook! This is your go to guide for Building with Llama: Getting started with Inference, Fine-Tuning, RAG. We also show you how to solve end to end problems using Llama model family and using them on various provider services

2

lmms-eval. Accelerating the development of large multimodal models (LMMs) with one-click evaluation module - lmms-eval.

2

Detic. Code release for "Detecting Twenty-thousand Classes using Image-level Supervision".

1

cvat_id_switch_solution. The repository contains the code to solve the id switches of tracks labelled using Intel's CVAT tool.

1

groundingLMM. Grounding Large Multimodal Model (GLaMM), the first-of-its-kind model capable of generating natural language responses that are seamlessly integrated with object segmentation masks [CVPR 2024].

1

easy-faster-rcnn.pytorch. An easy implementation of Faster R-CNN (https://arxiv.org/pdf/1506.01497.pdf) in PyTorch.

1

LLaVA-pp. 🔥🔥 LLaVA++: Extending LLaVA with Phi-3 and LLaMA-3 (LLaVA LLaMA-3, LLaVA Phi-3)

1

AutoGPT. An experimental open-source attempt to make GPT-4 fully autonomous.

1

multimodal-prompt-learning. Official repository of paper titled "MaPLe: Multi-modal Prompt Learning".

1

Video-ChatGPT. Python

1

DCL. Destruction and Construction Learning for Fine-grained Image Recognition

1

fairscale. PyTorch extensions for high performance and large scale training.

1

DisciplineViolationDetection. Discipline Anomaly Detection using Audio and Video Processing

1

LLaVA-pp-HF-Demo. Python

1

PyimagesearchComputerVisionCrashCourse. The source code for 17 days crash course by Pyimagesearch

1

EvalAI-Starters. How to create a challenge on EvalAI?

1

SwiftFormer. SwiftFormer: Efficient Additive Attention for Transformer-based Real-time Mobile Vision Applications

1

pytorch-stacked-hourglass. A PyTorch toolkit for 2D Human Pose Estimation.

1

detectron2. Detectron2 is FAIR's next-generation platform for object detection, segmentation and other visual recognition tasks.

1

object-centric-ovd. Official repository of paper titled "Bridging the Gap between Object and Image-level Representations for Open-Vocabulary Detection".

1

Restormer. [CVPR 2022--Oral] Restormer: Efficient Transformer for High-Resolution Image Restoration. SOTA for motion deblurring, image deraining, denoising (Gaussian/real data), and defocus deblurring.

1
32
Apply