Rare find

Qwen2.5-Omni. Qwen2.5-Omni is an end-to-end multimodal model by Qwen team at Alibaba Cloud, capable of understanding text, audio, vision, video, and performing real-time speech generation.

github.com/QwenLM/Qwen2.5-Omni

Vaya's read on this project

Problem, audience, market, and the verdict — sign in to see it.

Updates

No recent activity.