Mage-VL
Microsoft🇺🇸 USMajor lab· Jul 25, 2026
Mage-VL combines a from-scratch 4B visual encoder with a Qwen3-4B decoder to process images and video using codec-derived frame sparsity, cutting visual tokens by over 75%. Its dual-process design routes routine content through a lightweight gating mechanism while invoking the full model for event-worthy moments, enabling proactive streaming.
Summary of the model card
- License
- apache-2.0
- Modalities
- text, image → text
1 variant
| Standard | HF |
Sources
HuggingFaceBlogPaperDeutsch
Mage-VL kombiniert einen von Grund auf trainierten visuellen Encoder mit einem Qwen3-4B-Decoder und nutzt zur Verarbeitung von Bildern und Videos codec-abgeleitete Frame-Sparsamkeit, um die Anzahl der visuellen Token um über 75 Prozent zu reduzieren. Das Dual-Process-Design leitet Routineinhalte über eine einfache Gating-Funktion und ruft das Vollmodell bei ereignisrelevanten Momenten auf, wodurch proaktives Streaming moeglich ist.
Benchmarks
| Benchmark | Domain | Value | Evidence |
|---|---|---|---|
| Video-MMEvendor number | multimodal | 59.7 % correct | 81 |
| DocVQAvendor number | vision | 94.69 ANLS | 71 |
| AI2Dvendor number | vision | 81.54 % correct | 44 |
| ChartQAvendor number | vision | 83.96 % correct | 41 |
Vendor-reported numbers are marked and count with factor 0.7 in rankings. The evidence score says how much a benchmark still tells you today.
Provenance
- Detected
- Sep 27, 2026
- First source
- huggingface
- Description
- Summary of the model card
- Editorially reviewed
- —