This is your work, valued
Video-RAG-master. ✨✨[NeurIPS 2025] This is the official implementation of our paper "Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension"
★ 4493DRefTR. This is a PyTorch implementation of 3DRefTR proposed by our paper "A Unified Framework for 3D Point Cloud Visual Grounding"
★ 263DGCTR. This is a PyTorch implementation of 3DGCTR proposed by our paper “Rethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization”
★ 6PointMetaBase. This is a PyTorch implementation of PointMetaBase proposed by our paper "Meta Architecure for Point Cloud Analysis"
★ 1AI-Research-SKILLs. Comprehensive open-source library of AI research and engineering skills for any AI model. Package the skills and your claude code/codex/gemini agent will be an AI research agent with full horsepower. Maintained by Orchestra Research.
★ 11khumanizer. Agent skill that removes signs of AI-generated writing from text
★ 32kSupervisor-Skills. 将博导十年科研经验炼化为可直接调用的 AI 技能。从 Idea 构思到论文投稿,你的 AI 科研副导师。
★ 4.7kValley. Valley is a cutting-edge multimodal large model designed to handle a variety of tasks involving text, images, video, and audio data.
★ 290UltraEval-Audio. Your faithful, impartial partner for audio evaluation — know yourself, know your rivals. 真实评测,知己知彼。A unified benchmark framework for ASR/TTS/Audio Codec/audio LLM evaluation
★ 311AudioBench. AudioBench: A Universal Benchmark for Audio Large Language Models
★ 319MELD. MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversation
★ 1.1kMMAU. Python
★ 156paper-piggy. 🐷 论文小猪 PaperPiggy —— 让 AI agent 真正读懂你的文献库:39 个 MCP 工具 · 引文核对与论断三态核验 · 可追溯原文定位 · 可累积的综合层 Wiki。面向法学/社科研究者。
★ 1OmniScope. [ACMMM 2026🔥] This is the official implementation of our paper "OmniScope: Modality-decoupled Token Compression for Efficient Omnimodal Video Understanding"
★ 2academic-research-skills. Academic Research Skills for Claude Code: research → write → review → revise → finalize
★ 40kharness-engineering. Harness Engineering 学习指南 — 从概念理解到独立实践的深度学习档案
★ 5.1kfuture-agi. Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
★ 1.5kMindPipe. A powerful model compression framework for LLMs and LVLMs, adapted for NVIDIA GPUs and Huawei Ascend NPUs.
★ 1kPoCo. [CVPR2026] Official implementation of our paper “Rethinking Position Embedding as a Context Controller for Multi-Reference and Multi-Shot Video Generation”
★ 19Video-MME-v2. Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
★ 369SpecEyes. [ECCV 2026🔥] This is the official implementation of our paper "SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning"
★ 62WFS-SB. [CVPR 2026] Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
★ 32HunyuanOCR. HunyuanOCR-1.5: Making Lightweight OCR VLMs Faster and Better
★ 1.9kGOT-OCR2.0. Official code implementation of General OCR Theory: Towards OCR-2.0 via a Unified End-to-end Model
★ 8.2kSPEED. [ICLR'26] SPEED: Scalable, Precise, and Efficient Concept Erasure for Diffusion Models
★ 41SAFE. [KDD'25] Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective
★ 1122S-GDA. Python
★ 1Video-RAG-Ultra. Multi-modal video understanding system based on RAG technology
★ 2MLLM-Token-Compression. Towards Efficient Multimodal Large Language Models: A Survey on Token Compression
★ 214IPDN. Python
★ 173D-DRES. Python
★ 8Agent-Memory-Paper-List. The paper list of "Memory in the Age of AI Agents: A Survey"
★ 2.3kKTS. Kernel Temporal Segmentation
★ 64TimeChat. [CVPR 2024] TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
★ 425DDVC. This repo is a official codebase for our paper accepted to the ACL2025. The aim of this repo is to help other researchers
★ 7CAT-V. [AAAI 26 Demo] Offical repo for CAT-V - Caption Anything in Video: Object-centric Dense Video Captioning with Spatiotemporal Multimodal Prompting
★ 68LongCat-Flash-Omni. This is the official repo for the paper "LongCat-Flash-Omni Technical Report"
★ 500HunyuanVideo-1.5. HunyuanVideo-1.5: A leading lightweight video generation model
★ 4.5kOmniZip. [CVPR 2026] OmniZip: Audio-Guided Dynamic Token Compression for Fast Omnimodal Large Language Models
★ 103VideoMind. 🧠 VideoMind: A Chain-of-LoRA Agent for Temporal-Grounded Video Reasoning (ICLR 2026)
★ 350CLAP. Contrastive Language-Audio Pretraining
★ 2.2kMing. Ming - facilitating advanced multimodal understanding and generation capabilities built upon the Ling LLM.
★ 665STTM. [ICCV 2025] Multi-Granular Spatio-Temporal Token Merging for Training-Free Acceleration of Video LLMs
★ 61HoliTom. [NeurIPS 2025] HoliTom: Holistic Token Merging for Fast Video Large Language Models
★ 84VITA. The official implement of VITA, VITA15, LongVITA, VITA-Audio, VITA-VLA, and VITA-E.
★ 162Open-o3-Video. [ICML 2026] Official implementation of "Open-o3 Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence"
★ 158RWKV-LM. RWKV (pronounced RwaKuv) is an RNN with great LLM performance, which can also be directly trained like a GPT transformer (parallelizable). We are at RWKV-7 "Goose". So it's combining the best of RNN and transformer - great performance, linear time, constant space (no kv-cache), fast training, infinite ctx_len, and free sentence embedding.
★ 15kflash-linear-attention. 🚀 Efficient implementations for emerging model architectures
★ 5.5kDATE. Use 2 lines to empower absolute time awareness for Qwen2.5VL's MRoPE
★ 29VidCom2. [EMNLP 2025 Main] Video Compression Commander: Plug-and-Play Inference Acceleration for Video Large Language Models
★ 129VFlowOpt. [ICCV 2025] VFlowOpt: A Token Pruning Framework for LMMs with Visual Information Flow-Guided Optimization
★ 10ReAgent-V. [NeurIPS'25] ReAgent-V: A Reward-Driven Multi-Agent Framework for Video Understanding
★ 51mini-swe-agent. The 100 line AI agent that solves GitHub issues or helps you in your command line. Radically simple, no huge configs, no giant monorepo—but scores >74% on SWE-bench verified!
★ 6.1kAwesome-Multimodal-Token-Compression. [TMLR 2026] Survey: https://arxiv.org/pdf/2507.20198
★ 375Qwen3-Omni. Qwen3-omni is a natively end-to-end, omni-modal LLM developed by the Qwen team at Alibaba Cloud, capable of understanding text, audio, images, and video, as well as generating speech in real time.
★ 3.9kProject-Imaging-X. Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development
★ 468DHAT. [ICCV 2025] Towards Adversarial Robustness via Debiased High-Confidence Logit Alignment
★ 10LLaVA-OneVision-2. Fully Open Framework for Democratized Multimodal Training
★ 1.2kmcp_chatbot. A chatbot implementation compatible with MCP (terminal / streamlit supported)
★ 253sglang. SGLang is a high-performance serving framework for large language models and multimodal models.
★ 31kT2I-CoReBench. [ICLR'26] Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
★ 52minimind. 🧠「大模型」2小时完全从0训练64M的小参数LLM!Train a 64M-parameter LLM from scratch in just 2h!
★ 54kawesome-nano-banana. Awesome curated collection of images and prompts generated by gemini-2.5-flash-image (aka Nano Banana) state-of-the-art image generation and editing model. Explore AI generated visuals created with Gemini, showcasing Google’s advanced image generation capabilities.
★ 8.8kyoutu-agent. A simple yet powerful agent framework that delivers with open-source models
★ 4.6kViaRL. ViaRL: Adaptive Temporal Grounding via Visual Iterated Amplification Reinforcement Learning
★ 11ICoT. [CVPR' 25] Interleaved-Modal Chain-of-Thought
★ 112Awesome-MCoT. Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
★ 1kMINT-CoT. [NeurIPS 2025] MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
★ 107gpt-oss. gpt-oss-120b and gpt-oss-20b are two open-weight language models by OpenAI
★ 20kARC-Hunyuan-Video-7B. Structured Video Comprehension of Real-World Shorts
★ 240LLoVi. Official implementation for "A Simple LLM Framework for Long-Range Video Question-Answering"
★ 106VideoAgent. Python
★ 150BOLT. [CVPR2025] BOLT: Boost Large Vision-Language Model Without Training for Long-form Video Understanding
★ 55AKS. [CVPR 2025] Adaptive Keyframe Sampling for Long Video Understanding
★ 228unsloth-zoo. Utils for Unsloth https://github.com/unslothai/unsloth
★ 295verl. verl/HybridFlow: A Flexible and Efficient RL Post-Training Framework
★ 23kWan2.1. Wan: Open and Advanced Large-Scale Video Generative Models
★ 17kflash-attention. Fast and memory-efficient exact attention
★ 25kBackMix. [TPAMI2025] BackMix: Regularizing Open Set Recognition by Removing Underlying Fore-Background Priors
★ 16MEDAF. [AAAI2024] Exploring Diverse Representations for Open Set Recognition
★ 35Liger-Kernel. Efficient Triton Kernels for LLM Training
★ 6.5kMMPerspective. [NeurIPS 2025] A Comprehensive Benchmark for Perspective Perception, Reasoning, and Robustness
★ 13Do-You-See-Me. Python
★ 13RTV-Bench. [NeurIPS 2025] 𝓡𝓣𝓥-𝓑𝓮𝓷𝓬𝓱: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video.
★ 33ms-swift. Use PEFT or Full-parameter to CPT/SFT/DPO/GRPO 600+ LLMs (Qwen3.6, DeepSeek-V4, GLM-5.1, InternLM3, Llama4, ...) and 300+ MLLMs (Qwen3-VL, Qwen3-Omni, InternVL3.5, Ovis2.5, GLM4.5v, Gemma4, Llava, Phi4, ...) (AAAI 2025).
★ 15kVideoMathQA. VideoMathQA is a benchmark designed to evaluate mathematical reasoning in real-world educational videos
★ 24VidText. Comprehensive benchmark for video text understanding
★ 29OVO-Bench. [CVPR 2025] OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
★ 156SLOT. Python
★ 112Keye. Python
★ 809ROLL. An Efficient and User-Friendly Scaling Library for Reinforcement Learning with Large Language Models
★ 3.3kMM-EUREKA. MM-EUREKA: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
★ 771AdaVideoRAG. [NeurIPS 2025] AdaVideoRAG: Omni-Contextual Adaptive Retrieval-Augmented Efficient Long Video Understanding
★ 15VisualQuality-R1. [NeurIPS 2025 Spotlight] VisualQuality-R1 is the first open-sourced NR-IQA model can accurately describe and rate the image quality.
★ 201Awesome-Image-Quality-Assessment. A comprehensive collection of IQA papers
★ 1.5koss-browser. OSS Browser 提供类似windows资源管理器功能。用户可以很方便的浏览文件,上传下载文件,支持断点续传等。
★ 3.6kCo-Instruct. ④[ECCV 2024 Oral, Comparison among Multiple Images!] A study on open-ended multi-image quality comparison: a dataset, a model and a benchmark.
★ 87SILVR. Official Implementation for "SiLVR : A Simple Language-based Video Reasoning Framework"
★ 19AI-Agent-papers. Collection of recent works on AI Agents.
★ 17unsloth. Unsloth is a local UI for training and running Kimi K3, Gemma 4, Qwen3.6, DeepSeek, GLM and other models.
★ 69kLabel-Free-RLVR.
★ 311Temporal-R1. Reinforcement Learning Tuning for VideoLLMs: Reward Design and Data Efficiency
★ 62VAP. [NeurIPS 2025] Poison as Cure: Visual Noise for Mitigating Object Hallucinations in LVMs
★ 38Q-Insight. Q-Insight Family: Q-Insight, VQ-Insight and RALI (NeurIPS 2025 Spotlight, AAAI 2026 Oral, and ICLR 2026 Oral)
★ 314Video-Holmes. [ECCV 2026] Video-Holmes: Can MLLM Think Like Holmes for Complex Video Reasoning?
★ 95Omni-R1. [NeurIPS 2025] Official Repo of Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
★ 126LlamaFactory. Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
★ 74kPixel-Reasoner. Pixel-Level Reasoning Model trained with RL [NeuIPS25]
★ 301Paper-Writing-Tips. MLNLP社区用来帮助大家避免论文投稿小错误的整理仓库。 Paper Writing Tips
★ 4.6kSora-Generates-Videos-with-Stunning-Geometrical-Consistency. Sora Generates Videos with Stunning Geometrical Consistency
★ 51AIGCBench. [TBench 2024] Official implementation of "AIGCBench: Comprehensive Evaluation of Image-to-Video Content Generated by AI"
★ 48Minimal-RL. Python
★ 275KuaiMod.github.io. JavaScript
★ 68Qwen-VL-Series-Finetune. An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
★ 1.9kUnifiedReward. Official implementation of UnifiedReward & [NeurIPS 2025] UnifiedReward-Think & UnifiedReward-Flex
★ 796Q-Insight. Q-Insight is open-sourced at https://github.com/bytedance/Q-Insight. This repository will not receive further updates.
★ 142AnchorCrafter. Python
★ 669VITA-Audio. ✨✨[NeurIPS 2025] VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
★ 683trl. Train transformer language models with reinforcement learning.
★ 19kopen-r1. Fully open reproduction of DeepSeek-R1
★ 26kVideoChat-R1. [NIPS2025] VideoChat-R1 & R1.5: Enhancing Spatio-Temporal Perception and Reasoning via Reinforcement Fine-Tuning
★ 268HowToCook. Programmer's guide about how to cook at home.
★ 101kFRAG. Python
★ 15MR-Video. MR. Video: MapReduce is the Principle for Long Video Understanding
★ 31OpenRLHF. An Easy-to-use, Scalable and High-performance Agentic RL Framework based on Ray (PPO & DAPO & REINFORCE++ & VLM & TIS & vLLM & Ray & Async RL)
★ 9.9kVideo-MMLU. A Massive Multi-Discipline Lecture Understanding Benchmark
★ 34SkyReels-V2. SkyReels-V2: Infinite-length Film Generative model
★ 7.3kViLAMP. [ICML 2025] Official repository for paper "Scaling Video-Language Models to 10K Frames via Hierarchical Differential Distillation"
★ 195describe-anything. [ICCV 2025] Implementation for Describe Anything: Detailed Localized Image and Video Captioning
★ 1.5kVidCapBench. Python
★ 13tarsier. Tarsier -- a family of large-scale video-language models, which is designed to generate high-quality video descriptions , together with good capability of general video understanding.
★ 548VBench. [CVPR2024 Highlight] VBench - We Evaluate Video Generation
★ 1.7kVPO. Python
★ 25VisionReward. [AAAI 2026] VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation
★ 422HumanAesExpert. Official implementation of "HumanAesExpert: Advancing a Multi-Modality Foundation Model for Human Image Aesthetic Assessment"
★ 122LLaDA. Official PyTorch implementation for "Large Language Diffusion Models"
★ 3.9ktensorzero. TensorZero is an open-source LLMOps platform that unifies an LLM gateway, observability, evaluation, optimization, and experimentation.
★ 12kSymbolicDet. Python
★ 13VideoAlign. [NeurIPS 2025] Improving Video Generation with Human Feedback
★ 489vllm. A high-throughput and memory-efficient inference and serving engine for LLMs
★ 88kQwen2.5-Omni. Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.
★ 4.1kTime-R1. R1-like Video-LLM for Temporal Grounding
★ 138aurora. [ICLR 2025] AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
★ 147SlideChat. SlideChat: A Large Vision-Language Assistant for Whole-Slide Pathology Image Understanding
★ 126Campus2026. 2026届互联网校招&2025互联网实习信息汇总,欢迎共创
★ 411PE3R. [CVPR'26] PE3R: Perception-Efficient 3D Reconstruction. Take 2 - 3 photos with your phone, upload them, wait a few minutes, and then start exploring your 3D world via text!
★ 415Sparrow. Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
★ 32Vision-R1. [ICLR2026] This is the first paper to explore how to effectively use R1-like RL for MLLMs and introduce Vision-R1, a reasoning MLLM that leverages cold-start initialization and RL training to incentivize reasoning capability.
★ 1.6kE3-FaceNet. [ICML 2024] Fast Text-to-3D-Aware Face Generation and Manipulation via Direct Cross-modal Mapping and Geometric Regularization
★ 23StoryWeaver. [AAAI 2025] StoryWeaver: A Unified World Model for Knowledge-Enhanced Story Character Customization
★ 227ACL25-AdaReTaKe. Official implementation of paper AdaReTaKe: Adaptive Redundancy Reduction to Perceive Longer for Video-language Understanding
★ 91QuoTA. ✨✨[AAAI 2026] This is the official implementation of our paper "QuoTA: Query-oriented Token Assignment via CoT Query Decouple for Long Video Comprehension"
★ 79TPO. Python
★ 41Woodpecker. ✨✨Woodpecker: Hallucination Correction for Multimodal Large Language Models
★ 649LinVT. LinVT: Empower Your Image-level Large Language Model to Understand Videos
★ 83VideoLLaMA3. Frontier Multimodal Foundation Models for Image and Video Understanding
★ 1.2kVideoNIAH. VideoNIAH: A Flexible Synthetic Method for Benchmarking Video MLLMs
★ 57VILA. VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
★ 3.8kVideoChat-Flash. [ICLR2026] VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
★ 526Evolver. [COLING 2025🔥] Evolver: Chain-of-Evolution Prompting to Boost Large Multimodal Models for Hateful Meme Detection
★ 17LongVU. [ICML 2025] Official PyTorch implementation of LongVU
★ 431FrameFusion. [ICCV'25] The official code of paper "Combining Similarity and Importance for Video Token Reduction on Large Visual Language Models"
★ 75FastV. [ECCV 2024 Oral] Code for paper: An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
★ 591DyCoke. [CVPR 2025] DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
★ 114Awesome-Token-Compress. A paper list of some recent works about Token Compress for Vit and VLM
★ 944DeepSeek-R1.
★ 92kPopkart-Excess-system-modelling. Python
★ 3GaitRDAE. Official Repository of GaitRDAE.
★ 4LVBench. [ICCV 2025] LVBench: An Extreme Long Video Understanding Benchmark
★ 145NExT-QA. NExT-QA: Next Phase of Question-Answering to Explaining Temporal Actions (CVPR'21)
★ 189MiniCPM-V. A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
★ 26kvideo-ReTaKe. Official implementation of paper ReTaKe: Reducing Temporal and Knowledge Redundancy for Long Video Understanding
★ 40Thinking-Claude. Let your Claude able to think
★ 17kMBQ. The code repository of "MBQ: Modality-Balanced Quantization for Large Vision-Language Models"
★ 93ToMe. A method to increase the speed and lower the memory footprint of existing vision transformers.
★ 1.2kAIM. [ICCV 2025] Official code for "AIM: Adaptive Inference of Multi-Modal LLMs via Token Merging and Pruning"
★ 65TAR3D. Official Code for 'TAR3D: Creating High-Quality 3D Assets via Next-Part Prediction' (ICCV 2025)
★ 77InternVL. [CVPR 2024 Oral] InternVL Family: A Pioneering Open-Source Alternative to GPT-4o. 接近GPT-4o表现的开源多模态对话模型
★ 10kMMStar. [NeurIPS 2024] This repo contains evaluation code for the paper "Are We on the Right Way for Evaluating Large Vision-Language Models"
★ 215VLMEvalKit. Open-source evaluation toolkit of large multi-modality models (LMMs), support 220+ LMMs, 80+ benchmarks
★ 4.3kMDIN. [MM2024 Oral] 3D-GRES: Generalized 3D Referring Expression Segmentation
★ 43RG-SAN. [NeurIPS 2024 Oral] RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation
★ 20Sparrow. Repo for paper "T2Vid: Translating Long Text into Multi-Image is the Catalyst for Video-LLMs"
★ 48AutoGPTQ. An easy-to-use LLMs quantization package with user-friendly apis, based on GPTQ algorithm.
★ 5.1kbitsandbytes. Accessible large language models via k-bit quantization for PyTorch.
★ 8.4kSAVEn-Vid. SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context
★ 5QVLM. [NeurIPS'24]Efficient and accurate memory saving method towards W4A4 large multi-modal models.
★ 102video-rag.github.io. JavaScript
★ 1Video-RAG-master. ✨✨[NeurIPS 2025] This is the official implementation of our paper "Video-RAG: Visually-aligned Retrieval-Augmented Long Video Comprehension"
★ 449Freeze-Omni. ✨✨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
★ 388Video-XL. 🔥🔥First-ever hour scale video understanding models
★ 626LongVideoBench. [Neurips 24' D&B] Official Dataloader and Evaluation Scripts for LongVideoBench.
★ 134MLVU. 🔥🔥MLVU: Multi-task Long Video Understanding Benchmark
★ 266Qwen3-VL. Qwen3-VL is the multimodal large language model series developed by Qwen team, Alibaba Cloud.
★ 20kVideoAgent. This is the official code of VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding (ECCV 2024)
★ 321peiqianjichang. 赔钱机场官网地址
★ 847EgoSchema. Python
★ 117lmms-eval. One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
★ 4.3kDPR. Dense Passage Retriever - is a set of tools and models for open domain Q&A task.
★ 1.9k3DGCTR. This is a PyTorch implementation of 3DGCTR proposed by our paper “Rethinking 3D Dense Caption and Visual Grounding in A Unified Framework through Prompt-based Localization”
★ 6ircot. Repository for Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions, ACL23
★ 274mini-omni. open-source multimodal large language model that can hear, talk while thinking. Featuring real-time end-to-end speech input and streaming audio output conversational capabilities.
★ 3.6kYoLLaVA. 🌋👵🏻 Yo'LLaVA: Your Personalized Language and Vision Assistant (NeurIPS 2024)
★ 123video-mme.github.io. JavaScript
★ 2