LLaDA-UI
Ant Group (inclusionAI)๐จ๐ณ CNMajor labยท Sep 6, 2026
TextVisionopen weights
LLaDA-UI is an MoE-based, block-wise diffusion vision-language GUI agent. It understands screenshots at their native aspect ratio and produces grounded coordinates or structured actions for mobile, desktop, and web interfaces.
From the model card
- Modalities
- text, image โ text
1 variant
| Standard | HF |
Sources
HuggingFacePaperProvenance
- Detected
- Sep 27, 2026
- First source
- huggingface
- Description
- From the model card
- Editorially reviewed
- โ