ModelRadar
/

Ling-3.0-tiny

Ant Group (inclusionAI)🇨🇳 CNMajor lab· Aug 10, 2026

Textopen weights

Ling-3.0-tiny is a 7.9B hybrid MoE language model that activates only 1.3B parameters per token. It combines a 3:1 alternating layout of Kimi Delta Attention and Multi-Head Latent Attention layers with a sparse MoE feed-forward network of 128 experts (8 routed plus 1 shared per token). The model provides reasoning via a configurable thinking mode, supports multi-turn tool calling through XML-tagged function calls, and offers BF16, FP8, and INT4 quantized weights. Benchmarks report around 160+ tokens/s inference speed on an H20 GPU, approximately 8.3 GiB peak VRAM at 8K context in FP8, and an Artificial Analysis Intelligence Index score of 25 and Agentic Index score of 16. SGLang integration includes a built-in speculative decoding recipe (NEXTN/MTP) and YaRoN-based context extension up to 262K tokens. No explicit evaluation table values are published in the released material beyond those two index scores.

Summary of the model card

License
mit
Modalities
text → text

8 variants

SingprobeHF
GGUFHF
BaseHF
Base 30T30THF
Base MidtrainHF
INT4HF
FP8HF
StandardHF

Sources

HuggingFacePaper
Deutsch

Ling-3.0-tiny ist ein hybrides MoE-Sprachmodell mit 7,9 Milliarden Gesamtparametern, von denen pro Token nur 1,3 Milliarden aktiv sind. Die Architektur kombiniert abwechselnd Kimi Delta Attention und Multi-Head Latent Attention im Verhältnis 3 zu 1 sowie einen_sparse_ Feed-Forward-Bereich mit 128 Experten (8 geroutete plus 1 Shared Expert pro Token). Das Modell bietet reasoning via einem konfigurierbaren Thinking-Mode unterstuetzt mehrstufige Tool-Calls durch XML-formatierte Funktionsaufrufe und liefert Gewichte in BF16, FP8 und INT4. Benchmarkangaben nennen rund 160 Plus Token pro Sekunde Inferenzgeschwindigkeit auf einer H20-GPU, etwa 8,3 Gibyte Spitzen-VRAM bei 8K Kontextlaenge in FP8, einen Wert von 25 im Artificial Analysis Intelligence Index und 16 im Agentic Index. Die SGLang-Anbindung beinhaltet ein integriertes Speculative-Decoding-Rezept (NEXTN/MTP) und eine Kontexterweiterung ueber YaRoN bis zu 262K Tokens.

Provenance

Detected
Sep 27, 2026
First source
huggingface
Description
Summary of the model card
Editorially reviewed
—