This is your work, valued
Computer Vision, Multimodal AI
Qwen3-VL-Outpost. Qwen3-VL-Outpost is an experimental, high-performance visual reasoning and multimodal inference suite designed for advanced image analysis, optical character recognition, and complex scene understanding. Built around the state-of-the-art Qwen3-VL and Qwen2.5-VL model families.
9Text-Tokenizer-Playground. Text Tokenizer Playground ( Transformers.js ) SDK in Hugginface.
7Qwen3-VL-HF-Demo. The demo of Qwen3-VL-30B-A3B-Instruct, the next-generation and powerful vision-language model in the Qwen series, delivers comprehensive upgrades across the board — including superior text understanding and generation, deeper visual perception and reasoning, extended context length, enhanced spatial and video dynamics comprehensions.
6Medical-Term-Article-Search. HealthCare-Informatics-MediSearch
6Vision-Inference. What Happen Next ? Live Inference
6youtube-video-downloader. Enter YouTube link 🔗 To Download Video⬇️
6StrangerAI. Turning Ideas to Product - StrangerAI - StrangerZone. Recommended to Deploy inside Huggingface Spaces SDK as GRADIO
6Tiny-VLMs-Lab. Tiny VLMs Lab is a Hugging Face Space and open-source project showcasing lightweight Vision-Language Models for image captioning, OCR, reasoning, and multimodal understanding. It offers a simple Gradio interface to upload images, query models, adjust generation settings, and export results in Markdown or PDF.
6Bidirectional-and-Auto-Regressive-Transformer-CNN. BART’s primary task is used to generate clean semantically coherent text from corrupted text data but it can also be used for a variety of different NLP sub-tasks like language translation, question-answering tasks, text summarization, paraphrasing, etc.
6Airbnb-NYC-Maps. Airbnb Price in NYC ( Select Boroughs )
6EHRM-Demo. HTML
6Auto-Abliteration. modify a language model's behavior by abliterating its weights.
6Multimodal-OCR2. Multimodal-OCR2 is an advanced, experimental optical character recognition and document analysis suite designed to extract high-fidelity text, reconstruct complex document layouts, and generate structured markdown from diverse visual inputs.
5Photo-Mate-i2i. Photo-Mate-i2i – a space for experimenting with adapters for image manipulation using Kontext adapters, including Photo-Restore-i2i, PhotoCleanser-i2i, Polaroid-Warm-i2i, Yarn-Photo-i2i, Monochrome-Pencil.
5Vit-Mature-Content-Detection. Vit-Mature-Content-Detection is an image classification vision-language model fine-tuned from vit-base-patch16-224-in21k for a single-label classification task. It classifies images into various mature or neutral content categories using the ViTForImageClassification architecture.
5Master-GPT. Chat, Web, Media, Image GPT
5StableDiffusion. Continuous progress in AI research leads to the development of more robust algorithms, models, and techniques, making AI solutions more effective and reliable.
5PRITHIVSAKTHIUR.
5Photo-Background-Remover. Just remove background by a click
5Chatbot-GPT. 3-In-1-Chatbot - GPT
5Stable-Wallpapers. Demo space for generating, Desktop / Mobile Wallpapers. 16:9 / 9:16 #Dream Wallpaper by Stable Wallpaper [ stable diffusion xl ]
4Gemini-Image-Studio-HF. A state-of-the-art image generation and editing tool powered by Google's Generative AI models. This React-based web application allows users to generate images from text prompts, edit existing images, or create images from hand-drawn sketches.
4Qwen-Image-LoRA-DLC. Qwen-Image model with various LoRA (Low-Rank Adaptation) styles. This tool provides both speed-optimized and quality-focused generation modes with a curated collection of artistic styles.
4FLUX-LoRA-DLC2. FLUX-LoRA-DLC2 is an experimental, advanced image generation and image-to-image manipulation ecosystem. Built on top of the state-of-the-art black-forest-labs/FLUX.1-dev foundation, this application incorporates a dynamic multi-LoRA switching engine loaded with a comprehensive collection of over 100 stylistic adapters.
4Nano-Banana-AIO-HF. Nano Banana AIO HF is a web application built with React and the Google Gemini API for image generation and editing. It provides an all-in-one interface for creating, editing, and manipulating images using AI-powered tools.
4Flux-Image-Captioner. FLUX.1-dev with Qwen2VL Captioner and Prompt Enhancer
4Multimodal-VLM-Thinking. Demo for state-of-the-art Vision-Language Models (VLMs) for both image and video understanding tasks. This application offers a unified interface to interact with various specialized models for OCR, document analysis, visual reasoning, and multimodal understanding.
4Watermark-Detection-SigLIP2. Watermark-Detection-SigLIP2 is a vision-language encoder model fine-tuned from google/siglip2-base-patch16-224 for binary image classification. It is trained to detect whether an image contains a watermark or not, using the SiglipForImageClassification architecture.
4Doc-VLMs-exp. An experimental document-focused Vision-Language Model application that provides advanced document analysis, text extraction, and multimodal understanding capabilities. This application features a streamlined Gradio interface for processing both images and videos using state-of-the-art vision-language models specialized in document understanding.
4Banana-Zoom. Banana Zoom an advanced image enhancement web app that lets users select regions of an image for AI-powered upscaling and detail refinement. Using Google’s (nano banana)
4Gen-Vision. Multiple Conditioned Image Generation, SDXL, Low-rank adaptation Refined
4VLM-Parsing. VLM-Parsing is a Gradio-based web application for parsing documents and images into structured HTML and Markdown formats using advanced Vision Language Models (VLMs).
4VLM-Video-Understanding. A minimalistic demo for image inference and video understanding using OpenCV with some popular open-source VLMs
4Qwen-Image-Diffusion. Generate high-quality images from text prompts using the Qwen diffusion model with an intuitive Gradio web interface.
3Age-Classification-SigLIP2. Age-Classification-SigLIP2 is an image classification vision-language encoder model fine-tuned from google/siglip2-base-patch16-224 for a single-label classification task. It is designed to predict the age group of a person from an image using the SiglipForImageClassification architecture.
3Core-OCR. Core-OCR is an advanced, experimental Optical Character Recognition (OCR) and document analysis suite designed for highly accurate text extraction, table reconstruction, and complex visual reasoning. Built on the robust Qwen2.5-VL and Qwen2-VL multimodal architectures.
3GLM-4.1V-9B-Thinking-Video-Understanding. GLM-4.1V-9B-Thinking, designed to explore the upper limits of reasoning in vision-language models. By introducing a "thinking paradigm" and leveraging reinforcement learning, the model significantly enhances its capabilities.
3AIorNot-SigLIP2. AIorNot-SigLIP2 is a vision-language encoder model fine-tuned from google/siglip2-base-patch16-224 for binary image classification. It is trained to detect whether an image is generated by AI or is a real photograph using the SiglipForImageClassification architecture.
3Text-to-Image. Text to Image Gen [ Demo ]
3Gemini-Image-Studio. A state-of-the-art image generation and editing tool powered by Google's Generative AI models. This React-based web application allows users to generate images from text prompts, edit existing images, or create images from hand-drawn sketches.
3Qwen2.5-VL-7B-Instruct-Demo. Qwen2_5_VLForConditionalGeneration
3Aya-Vision-Ocr-vs-Qwen2VL-Ocr. Messy Handwriting OCR Comparison Between Aya-Vision-8B and Qwen2VL-OCR-2B
3Augmented-Waste-Classifier-SigLIP2. Augmented-Waste-Classifier-SigLIP2 is an image classification vision-language encoder model fine-tuned from google/siglip2-base-patch16-224
3VisionScope-R2. VisionScope-R2 is an experimental, highly versatile vision suite designed for advanced image inference, spatial reasoning, and complex scene understanding. Built upon the powerful Qwen2.5-VL and Qwen2-VL architectures.
3Human-Action-Recognition. Human-Action-Recognition is an image classification vision-language encoder model fine-tuned from google/siglip2-base-patch16-224 for multi-class human action recognition. It uses the SiglipForImageClassification architecture to predict human activities from still images.
3YOLOX-CPU. Ultralytics, YOLO v8 - Computer Vision
3POINTS-Reader-OCR. The POINTS-Reader, a vision-language model for end-to-end document conversion, is a powerful, distillation-free Vision-Language Model that sets new SoTA benchmarks.
3Yolo-NMS-Captioning. Object Detection - Captioning ( yolo8n & blip-image-captioning-large )
3Client-Record-CURD-OPs-Exercise. Client Record Management - CURD OPs + Blazor Web Assembly with Standalone App
3YOLOX-T4. Ultralytics, YOLO v8 - Computer Vision
3QwQ-Edge. All in Chatbot
3Doc-VLMs-v2-Localization. Doc-VLMs-v2-Localization is a demo app for the Camel-Doc-OCR-062825 model, fine-tuned from Qwen2.5-VL-7B-Instruct for advanced document retrieval, extraction, and analysis. It enhances document understanding and also integrates other notable Hugging Face models.
3SigLIP2-MultiDomain-App. SigLIP2 is a vision-language encoder model fine-tuned from google/siglip2-base-patch16-224
3Omni-Reasoner-Vision. Omni Reasoner for Vision
3Multimodal-VLMs. A comprehensive Gradio-based interface for running multiple state-of-the-art Vision-Language Models (VLMs) for Optical Character Recognition (OCR) and Visual Question Answering (VQA) tasks.
2PhotoCleanser-i2i. PhotoCleanser-i2i is an adapter for black-forest-lab's FLUX.1-Kontext-dev. It is an experimental LoRA designed for removing specified object(s) while preserving the remaining content in the image.
2FLUX.1-Comparator-Krea-Dev. A high-performance Gradio application for comparing and generating images using two powerful FLUX.1 diffusion models: FLUX.1-dev-merged and FLUX.1-krea-merged-dev. This application provides an intuitive interface for AI-powered image generation with advanced customization options.
2tooth-agenesis-siglip2. tooth-agenesis-siglip2 is a vision-language encoder model fine-tuned from google/siglip2-base-patch16-512 for multi-class image classification. It is trained to detect various dental anomalies and conditions such as Calculus, Caries, Gingivitis, Mouth Ulcer, Tooth Discoloration, and Hypodontia.
2Nvidia-Cosmos-Reason1-Demo. Physical AI models understand physical common sense and generate appropriate embodied decisions in natural language through long chain-of-thought reasoning processes.
2Imgscope-OCR-2B-0527. The Imgscope-OCR-2B-0527 model is a fine-tuned version of Qwen2-VL-2B-Instruct, specifically optimized for messy handwriting recognition, document OCR, realistic handwritten OCR, and math problem solving with LaTeX formatting. This model is trained on custom datasets for document and handwriting OCR tasks and textual understanding
2