This is your work, valued
Mega-ASR. First foundation ASR built for the real world - 7 atomic acoustic conditions, 54 compound scenarios, 2.6M samples, and up to ~30% gains over SOTA where every other model falls apart. **You'll come back to MEGA-ASR, after the rest fail in the wild. ⭐**
1.1kAudio-Interaction. Python
574Audio-Reasoner. The first Large Audio Language Model that enables native in-depth thinking, which is trained on large-scale audio Chain-of-Thought data.
297Pask. Towards Self-Evolving Proactive AI with Perpetual Memory
208Mini-Omni-Reasoner. Mini-Omni-Reasoner: a real-time speech reasoning framework that interleaves silent reasoning tokens with spoken response tokens (“thinking-in-speaking”), exploiting the LLM–audio throughput gap to keep speech fluent and low-latency while maintaining structured internal reasoning.
166Voices-in-the-Wild-Bench. Python
27