Rare find

ComfyUI_VLM_nodes. ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via llama.cpp, and hosted LLM/VLM APIs.

github.com/gokayfem/ComfyUI_VLM_nodes

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.