VibeVoice-ASR-Streaming
Microsoft🇺🇸Major lab
VibeVoice-ASR-Streaming is a unified streaming ASR model that transcribes Who (Speaker) said What (Content), with support for Customized Hotwords and 10 languages.
Releases from frontier labs, major companies and research institutes — grouped by model, deduplicated across HuggingFace, OpenRouter and official blogs.
Click a week to filter
Microsoft🇺🇸Major lab
VibeVoice-ASR-Streaming is a unified streaming ASR model that transcribes Who (Speaker) said What (Content), with support for Customized Hotwords and 10 languages.
Microsoft🇺🇸Major lab
Mage-ViT is the visual encoder at the core of Mage-VL. It is a Codec-ViT built primarily for video, where a single image is simply the degenerate one-frame case.
Microsoft🇺🇸Major lab
Mage-VL combines a from-scratch 4B visual encoder with a Qwen3-4B decoder to process images and video using codec-derived frame sparsity, cutting visual tokens by over 75%. Its dual-process design routes routine content through a lightweight gating mechanism while invoking the full model for event-worthy moments, enabling proactive streaming.
Microsoft🇺🇸Major lab
VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required.
Microsoft🇺🇸Major lab
Microsoft's specialized code-search model for coding agents — powers the 'Explore' subagent in SWE-FastContext. Not a general-purpose LLM, but relevant for agentic setups.