ModelRadar
/

What shipped in AI.

Releases from frontier labs, major companies and research institutes — grouped by model, deduplicated across HuggingFace, OpenRouter and official blogs.

Releases per weekFrontierMajor labResearch
10Jun 15Jul 13Aug 10Sep 7Sep 28

Click a week to filter

Last week20 models

MiMo-V2.6-Flash-MOPD

Xiaomi MiMo🇨🇳Major lab

Textopen

MiMo-V2.6-Flash-MOPD is an open-weight multimodal model by Xiaomi supporting text, images, audio, and video with tool-use capability and structured reasoning via <think> tags. Built on a 309B-parameter sparse MoE with 15B activated tokens per step and 1M token context, it addresses tool-call repetition from its predecessor through MOPD distillation using domain-specialized teachers.

Same release

Gemini 3.8 Live

Google DeepMind🇺🇸Frontier

Text

Gemini 3.8 Live erweitert ein Live-Dialytemodell um eine Videoavatar-Funktion, die Sprache und visuelles Feedback simultan verarbeitet — mit Lippen-Sync, multi-modalen Eingaben (Video plus Audio) und asynchronem Tool-Calling waehrend des Dialoges. Es laeuft nahtlos ueber 97 Sprachen hinweg ohne Abbruch der Videofidelitaet.

GLM 5.3 Prime

Z.ai (Zhipu)🇨🇳Frontier

TextReasoningAPI

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. It supports text input and output with a 1M-token...

Qwen3.8 Max Prime

Alibaba Qwen🇨🇳Frontier

TextReasoningVisionAPI

Qwen3.8 is a causal language model with vision capabilities, designed for complex tasks involving coding, professional work, and research. It supports a context length of up to 1,000,000 tokens and is compatible with various deployment tools.

Aion 3.5 Mini

AionLabs🇮🇱Major lab

TextReasoningAPI

Aion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, lower-cost sibling of Aion 3.5 and uses...

Same release
⧉ 262K$0.7 / $1.40OR

Solar Mini 4

Upstage🇰🇷Major lab

TextReasoningAPI

Solar 4 is a model family by Upstage. The Pro variant supports a context window of 512K tokens and is available via OpenRouter under the ID upstage/solar-pro4.

⧉ 524K$0.05 / $0.2OR

NV-Reason-CT Open 3D CT VLM

NVIDIA🇺🇸Major lab

TextVisionopen

NV-Reason-CT processes 3D CT volumes natively (not slice-by-slice) and generates structured diagnostic reports with step-by-step chain-of-thought reasoning validated by NIH radiologists. It scores state-of-the-art on CT-RATE (Macro-F1 0.614).

Space Bunny Alpha

◌ Stealth

Stealth🌐Stealth

TextReasoningVisionAPI

Space Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support. It delivers adjustable reasoning effort, and a 1M-token context window.

1 variantAlpha

GPT-6 Luna Pro

OpenAI🇺🇸Frontier

TextReasoningVisionAPI

GPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoningreasoning-mode

1 variantBatch
Same release

Claude Opus 5.5

Anthropic🇺🇸Frontier

TextReasoningVisionAPI

Anthropic's flagship model with significantly improved price–performance ratio: it communicates more directly than its predecessor, matches Fable 5.1's intelligence at lower cost, and achieves higher token efficiency, making it suitable for extended reasoning and agentic workflows.

1 variantBatch

Grok 4.7

xAI🇺🇸Frontier

TextReasoningVisionAPI

Anthropic's Opus-tier flagship addresses prior feedback by communicating clearly; delivers Fable 5.1 intelligence at a 20% lower price and is notably token-efficient, especially for long agentic workflows via slashed cache prices.

Qwen3.8 Omni Flash

Alibaba Qwen🇨🇳Frontier

TextReasoningVisionAPI

Qwen3.8 is a causal language model with vision capabilities, designed for complex tasks involving coding, professional work, and research. It supports a context length of up to 1,000,000 tokens and is compatible with various deployment tools.

MiMo-V2.6-Pro-UltraSpeed

Xiaomi MiMo🇨🇳Major lab

TextReasoningVisionAPI

MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it matches the original model in quality while delivering roughly 10x...

MiMo-V2.6-Flash

Xiaomi MiMo🇨🇳Major lab

TextReasoningVisionopenAPI

A model variant within the MiMo-V2.6 family using a sparse Mixture-of-Experts architecture with 15 billion active parameters out of 309 billion total. Processes text, images, video, and audio in one model. Trained with groupwise reinforcement learning designed to expand capabilities through self-improvement — targeting complex agentic tasks, tool use, and multi-session long contexts.

1 variantRL
Same release

Week of September 14, 20265 models

GLM 5.3 FlashX

Z.ai (Zhipu)🇨🇳Frontier

TextReasoningVisionAPI

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...

Ming-Image-0.1-Design

Ant Group (inclusionAI)🇨🇳Major lab

Imageopen

Ming-Image-0.1-Design is a 6B text-to-image model for UI, infographics, posters, and other text-rich visual designs. It generates complete visual compositions and supports RGBA output with transparent backgrounds.

Ming-Image-0.1-Design-Layer

Ant Group (inclusionAI)🇨🇳Major lab

Otheropen

Ming-Image-0.1-Design-Layer decomposes a flattened design image into a requested number of RGBA layers using an image and a layer plan.

Realtime-Venus

Ant Group (inclusionAI)🇨🇳Major lab

TextVisionopen

A full-duplex interaction system with asynchronous delegation

Qwen-Image-2.1

Alibaba Qwen🇨🇳Frontier

Imageopen

Qwen-Image-2.1 is a 7B-parameter image-generation and editing model supporting native RGBA output, multi-reference edits (up to 10), and resolutions from 1536x2752 up to 2752x1536.

2 variantsPE I2iPE T2i

Week of September 7, 20266 models

Fugu Ultra v2

Sakana AI🇯🇵Major lab

TextReasoningVisionAPI

Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...

Fugu Max

Sakana AI🇯🇵Major lab

TextReasoningVisionAPI

Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...

Atria-Dawn

Shanghai AI Lab (InternLM)🇨🇳Research

Otheropen

Atria Dawn Preview is a preview release of a new-generation agentic model developed by the Shanghai Artificial Intelligence Laboratory.

3 variantsPreview Ascend W8A8Preview FP8Preview

DeepSeek V4.1 Flash

DeepSeek🇨🇳Frontier

TextReasoningVisionopenAPI

Multimodal MoE model that natively processes images and text with a 1M-token context window. Based on a 552B backbone, it activates 8B/16B per step, cutting KV-cache overhead to roughly a quarter of its predecessor.

1 variantBatch

ChatGPT Images 2.5

OpenAI🇺🇸Frontier

Image

ChatGPT Images 2.5 helps turn your ideas, sketches, and reference photos into more personalized, polished images that better reflect your ideas.

Mercury 2.5

Inception🇺🇸Major lab

TextReasoningAPI

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

⧉ 260K$0.04 / $0.15OR

Week of August 31, 202616 models

LLaDA-UI

Ant Group (inclusionAI)🇨🇳Major lab

TextVisionopen

LLaDA-UI is an MoE-based, block-wise diffusion vision-language GUI agent. It understands screenshots at their native aspect ratio and produces grounded coordinates or structured actions for mobile, desktop, and web interfaces.

LLaDA2.2-mini

Ant Group (inclusionAI)🇨🇳Major lab

Textopen

LLaDA2.2-mini is the lightweight variant of the agentic diffusion language model in the LLaDA2 series.

GPT-6 Astra

OpenAI🇺🇸Frontier

TextReasoningVisionAPI

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

1 variantBatch
Same release

Ling 3.0 Flash VL

Ant Group (inclusionAI)🇨🇳Major lab

TextReasoningVisionopenAPI

Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...

3 variantsINT4FP4FP8

Ling 3.0 Flash Sante

Ant Group (inclusionAI)🇨🇳Major lab

TextReasoningAPI

Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for...

1 variantFree

EVIE

Tencent🇨🇳Major lab

Otheropen

EVIE-8B ist ein vision-language Modell für visuelle Dokumentensuche und -abfrage basierend auf Qwen3.5 mit bidirektionaler Selbstachtund Multi-Vector-Embeddings. Es dient als Distillierungslehrer für die kompakte EVIE-4.5B-Variante.

2 variants8B4.5B

Qwen3.8 Max

Alibaba Qwen🇨🇳Frontier

TextReasoningVisionAPI

Qwen3.8 is a causal language model with vision capabilities, designed for complex tasks involving coding, professional work, and research. It supports a context length of up to 1,000,000 tokens and is compatible with various deployment tools.

WeatherNext 3

Google DeepMind🇺🇸Frontier

Other

AI weather model that uses real-time satellite data instead of traditional physics simulations to produce hourly forecasts at five times the resolution of its predecessor. Also tracks precipitation and includes clean-energy variables.

Muse Spark 1.3

Meta🇺🇸Frontier

TextReasoningVisionAPI

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

1 variantContributor

Gemini 3.8 Flash

Google DeepMind🇺🇸Frontier

TextReasoningVisionAPI

Gemini 3.8 Live adds near real-time streaming video with expressive avatars — it processes audio and visual inputs simultaneously, supports precise lip-syncing across 97 languages, and can execute tool calls asynchronously while maintaining an uninterrupted conversation.

1 variantBatch

VibeVoice-ASR-Streaming

Microsoft🇺🇸Major lab

Audioopen

VibeVoice-ASR-Streaming is a unified streaming ASR model that transcribes Who (Speaker) said What (Content), with support for Customized Hotwords and 10 languages.

2 variants1.5B7B

Tiny Aya En Thinker

Cohere🇨🇦Major lab

Textopen

No description yet — the model card or announcement hasn't been summarised.

Tiny Aya L2 Thinker

Cohere🇨🇦Major lab

Textopen

No description yet — the model card or announcement hasn't been summarised.

Claude Fable 5.1

Anthropic🇺🇸Frontier

TextReasoningVisionAPI

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

1 variantBatch

Jina OCR V1

Jina AI🇩🇪Major lab

TextVisionopen

jina-ocr-v1 is an end-to-end document parsing model designed for high-quality OCR at an efficient serving point.

Week of August 24, 20269 models

LLaDA-Image-Turbo

Ant Group (inclusionAI)🇨🇳Major lab

Imageopen

LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes

1 variantFP8
Same release

Drive-1.0

Alibaba Qwen🇨🇳Frontier

TextVisionopen

Qwen-Drive-1.0 retains the architecture of the pretrained Qwen3.5 vision-language model and integrates 3D perception, visual question answering, and motion planning within a unified framework.

1 variant4B

Hy4

Tencent🇨🇳Major lab

TextReasoningopenAPI

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that...

3 variantsPreviewPreview FP8Preview

Ling 3.0 Flash Fin

Ant Group (inclusionAI)🇨🇳Major lab

TextReasoningopenAPI

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...

4 variantsFreeFP4INT4FP8

MiniCPM5

OpenBMB🇨🇳Research

Textopen

MiniCPM5-2B ist ein dichter 2B-Transformer von OpenBMB, optimiert für On-Device-Einsatz und ressourcenbeschränkte Szenarien. Das Modell konkurrenzieren 4B-Klassenmodelle und zeigt Vorteile in Code, Mathematik, Langkontext-Verständnis sowie Tool-Use und agentic Tasks.

9 variants2B2B Dspark GGUF2B GPTQ2B Dspark2B GGUF

GLM 5.3 Flash

Z.ai (Zhipu)🇨🇳Frontier

TextReasoningVisionopenAPI

GLM 5.3 is a language model developed by Z.AI (Zhipu) with a large context window of 1,048,576 tokens.

2 variantsBatchBF16

Qwen3.8 Flash

Alibaba Qwen🇨🇳Frontier

TextReasoningVisionopenAPI

Qwen3.8 is a causal language model with vision capabilities, designed for complex tasks involving coding, professional work, and research. It supports a context length of up to 1,000,000 tokens and is compatible with various deployment tools.

1 variantFP8

Nemotron-3-Diarization

NVIDIA🇺🇸Major lab

Otheropen

Nemotron 3 Diarization is an open-weight speaker diarization model designed to determine "who spoke when" in real-world audio. It supports both streaming and offline inference and handles up to eight speakers.

1 variantPreview

Week of August 17, 20264 models

DeepSeek V4 Flash Vision

DeepSeek🇨🇳Frontier

TextReasoningVisionopenAPI

Multimodal text-and-image model with substantially improved agent capabilities versus its text-only predecessor; competes in text-only agent tasks but trails leading proprietary models.

2 variantsExpExp

AuK-Flash

Tencent🇨🇳Major lab

Audioopen

- [2026/09/09] 🎉 AuK is now open-source. Code and model weights are publicly available.

Same release

GLM 5.3

Z.ai (Zhipu)🇨🇳Frontier

TextReasoningopenAPI

GLM 5.3 is a language model developed by Z.AI (Zhipu) with a large context window of 1,048,576 tokens.

2 variantsBatchBF16

Week of August 10, 202611 models

North-Small-Translate-1.0

Cohere🇨🇦Major lab

Textopen

No description yet — the model card or announcement hasn't been summarised.

2 variantsW4A16FP8

Gemini 3.7 Flash

Google DeepMind🇺🇸Frontier

TextReasoningVisionAPI

Gemini 3.7 is a model developed by Google DeepMind, featuring two variants: Flash and Flash Batch. It has a context window of 1,048,576 tokens.

1 variantBatch

Cmd

NVIDIA🇺🇸Major lab

Videoopen

Hmrishav Bandyopadhyay 1,2 , Xuanchi Ren 1 , Zijian Huang 1 , Jay Zhangjie Wu 1 , Tianshi Cao 1 , Ruilong Li 1 , Bryan Chu 1 , Sanja Fidler 1 , Yi-Zhe Song 2 , Zian Wang 1

Grok 4.6

xAI🇺🇸Frontier

TextReasoningVisionAPI

Grok 4.6 is a model developed by xAI, featuring a context window of 500,000 tokens. It is available in a standard variant.

Seed 2.1 Turbo

ByteDance Seed🇨🇳Major lab

TextReasoningVisionAPI

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...

⧉ 262K$0.5 / $2.50OR

Seed-2.0-Code

ByteDance Seed🇨🇳Major lab

TextReasoningVisionCodeAPI

Seed-2.0 is a model developed by ByteDance Seed, designed for code-related tasks with a large context window.

Namazu

Sakana AI🇯🇵Major lab

TextReasoningVisionAPI

Namazu is a model from Sakana AI with a 256K token context window (262144). No information is available on parameter count, licensing, weight openness, intended use cases, or capability tier.

Solar Pro 4

Upstage🇰🇷Major lab

TextReasoningAPI

Solar 4 is a model family by Upstage. The Pro variant supports a context window of 512K tokens and is available via OpenRouter under the ID upstage/solar-pro4.

⧉ 524K$0.09 / $0.36OR

North-Micro-Vision

Cohere🇨🇦Major lab

TextVisionopen

North-Micro-Vision is a 2.4B open-weight vision-language model combining a 400M native-resolution vision encoder with a 2B Command A+ language backbone, designed for prototyping and fine-tuning.

1 variantInstruct

Ling-3.0-tiny

Ant Group (inclusionAI)🇨🇳Major lab

Textopen

Ling-3.0-tiny is a 7.9B hybrid MoE language model that activates only 1.3B parameters per token. It combines a 3:1 alternating layout of Kimi Delta Attention and Multi-Head Latent Attention layers with a sparse MoE feed-forward network of 128 experts (8 routed plus 1 shared per token). The model provides reasoning via a configurable thinking mode, supports multi-turn tool calling through XML-tagged function calls, and offers BF16, FP8, and INT4 quantized weights. Benchmarks report around 160+ tokens/s inference speed on an H20 GPU, approximately 8.3 GiB peak VRAM at 8K context in FP8, and an Artificial Analysis Intelligence Index score of 25 and Agentic Index score of 16. SGLang integration includes a built-in speculative decoding recipe (NEXTN/MTP) and YaRoN-based context extension up to 262K tokens. No explicit evaluation table values are published in the released material beyond those two index scores.

7 variantsSingprobeGGUFBaseBase 30TBase Midtrain

LTX-2.5-Pre-Trained

Lightricks🇮🇱Major lab

Videoopen

No description yet — the model card or announcement hasn't been summarised.

Week of August 3, 20267 models

Muse Glimmer

Meta🇺🇸Frontier

TextReasoningVisionopenAPI

Muse Glimmer is a 30-billion-parameter causal language model designed for local agentic AI workflows, optimized for NVIDIA hardware. It integrates multi-step reasoning, reliable tool use, and multimodal understanding, enabling complex task completion without cloud dependency.

5 variants30B30B Executorch Pte30B Assistant30B GGUF30B

Granite 4.2

IBM🇺🇸Major lab

TextReasoningopenAPI

Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning.

19 variants8B30B BF16 MLX8B BF16 MLX3B BF16 MLX8B NVFP4

Music3

MiniMax🇨🇳Major lab

Audioopen

MiniMax Music 3 is a music generation model capable of creating complete songs up to five minutes long from lyrics and a style description. It accepts lyrics with section tags such as verse, chorus, bridge, and outro, along with structured metadata covering genre, tempo, instrumentation, and vocal characteristics. The model produces stereo audio at 32 kHz and maintains thematic coherence across full song structures including intro, development, and finale.

LFM2.5-VL

Liquid AI🇺🇸Major lab

TextVisionopen

LFM2.5-VL-3B is a multimodal variant of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on LFM2-VL-3B with further mid- and post-training.

10 variants3B Dspark GGUF3B Dspark3B GGUF3B3B MLX BF16

Muse Spark 1.2

Meta🇺🇸Frontier

TextReasoningVisionAPI

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...

1 variantContributor

Qwen3.8

Alibaba Qwen🇨🇳Frontier

TextReasoningVisionopenAPI

Qwen3.8 is a causal language model with vision capabilities, designed for complex tasks involving coding, professional work, and research. It supports a context length of up to 1,000,000 tokens and is compatible with various deployment tools.

7 variants27B27B Free2.4T A95B27B27B FP8

Granite Speech 5.0

IBM🇺🇸Major lab

Audioopen

Model Summary: Granite Speech 5.0 TurboCTC Non-commercial is a compact 470 million parameter English ASR model with very high inference speed that is released for research and noncommercial purposes only. We invite users to refer to granite-speech-5.0-470m-turboctc for other use cases.

2 variants470M Turboctc Nc470M Turboctc

Week of July 27, 20269 models

Nemotron 3.5 Lightning

NVIDIA🇺🇸Major lab

TextReasoningopenAPI

Nemotron 3.5 Lightning ist ein Textmodell von NVIDIA mit einem Kontextfenster von 262 K Tokens. Es ist seit dem 11. August 2026 auf OpenRouter verfuegbar.

6 variantsFree30B A3B Base BF1630B A3B NVFP4 Dspark30B A3B NVFP4 Dflash30B A3B NVFP4

LongCat-Flash-Lite-Sparse

Meituan LongCat🇨🇳Major lab

Textopen

LongCat-Flash-Lite-Sparse is a non-thinking Mixture-of-Experts (MoE) model with 69B total parameters and approximately 3B activated parameters per token. Built on LongCat-Flash-Lite, it replaces dense MLA with LongCat Sparse Attention (LSA) and natively supports context lengths of up to 1M tokens.

NemotronLabs-VoiceChat

NVIDIA🇺🇸Major lab

Otheropen

▶ Hear it first. Natural turn-taking, barge-in and live tool calling.

1 variant11B

K-EXAONE-2.0

LG AI Research🇰🇷Major lab

Textopen

We introduce K-EXAONE 2.0, a frontier-scale multilingual language model developed by LG AI Research. K-EXAONE 2.0 was scaled to more than three times the size of its predecessor through upcycling, followed by continual pretraining, difficulty-focused mid-training, and post-training.

4 variants750B A37B Dspark750B A37B NVFP4750B A37B FP8750B A37B

Intern-S2-Mobius

Shanghai AI Lab (InternLM)🇨🇳Research

TextVisionopen

We introduce Intern-S2-Mobius, a 35B foundation model built on the Mobius-v0 architecture realized by Xtuner and LMDeploy.

1 variantFP8

UEmbed

Alibaba Tongyi Lab🇨🇳Research

Embeddingopen

UEmbed is a decoder-only multimodal embedding model that produces both dense embeddings and SPLADE-style sparse lexical embeddings from a single causal forward pass. It supports text, image, video, and mixed-modal inputs for retrieval, multimodal search, and visual-document retrieval.

3 variants9B4B2B

MiniMax-H3

MiniMax🇨🇳Major lab

Videoopen

MiniMax H3 is a general-purpose video generation model with native stereo audio support. It takes text, images, videos, and audio as inputs and outputs synchronized audio-video clips between 4 and 15 seconds long at up to 2K resolution and 24 FPS. The input side is highly flexible, supporting first-and-last-frame mode, multi-reference inputs with up to nine images, three video clips, and three audio clips combined. A dedicated preprocessing component called H3-Context-IR parses and relates the multimodal context before generation. The model communicates fluently in eleven languages. Performance details beyond the specification sheet are not covered in the provided sources.

Qwen3.7 Flash

Alibaba Qwen🇨🇳Frontier

TextReasoningVisionAPI

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

Inkling Small

Thinking Machines🇺🇸Major lab

TextReasoningVisionopenAPI

Inkling Small is a general-purpose multimodal model accepting text, image, and audio inputs to generate text. Its 276B-parameter sparse MoE architecture uses 12B active parameters per token, making it suited for agentic systems, coding assistance, and conversation.

2 variantsFreeNVFP4

Week of July 20, 202615 models

Mage-ViT

Microsoft🇺🇸Major lab

Embeddingopen

Mage-ViT is the visual encoder at the core of Mage-VL. It is a Codec-ViT built primarily for video, where a single image is simply the degenerate one-frame case.

Mage-VL

Microsoft🇺🇸Major lab

TextVisionopen

Mage-VL combines a from-scratch 4B visual encoder with a Qwen3-4B decoder to process images and video using codec-derived frame sparsity, cutting visual tokens by over 75%. Its dual-process design routes routine content through a lightweight gating mechanism while invoking the full model for event-worthy moments, enabling proactive streaming.

Claude Opus 5

Anthropic🇺🇸Frontier

TextReasoningVisionAPI

Claude Opus 5.5 improves communication clarity over previous Opus models and achieves token efficiency approaching Fable 5.1 at lower cost. Output is capped at 128K tokens, making extended chain-of-thought unreliable.

1 variantBatch

VibeVoice-ASR-BitNet

Microsoft🇺🇸Major lab

Audioopen

VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required.

Apertus-v1.5

Swiss AI Initiative🇨🇭Research

TextVisionopen

No description yet — the model card or announcement hasn't been summarised.

2 variants70B8B

Hy-MT2

Tencent🇨🇳Major lab

TextopenAPI

Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided...

4 variants1.8B30B A3B7B30B A3B GGUF

Ling 3.0 Flash

Ant Group (inclusionAI)🇨🇳Major lab

TextReasoningopenAPI

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers...

8 variantsSingprobeGGUFBaseBase 30TBase Midtrain

LTX-2.5

Lightricks🇮🇱Major lab

Videoopen

No description yet — the model card or announcement hasn't been summarised.

1 variantDiffusers

KAT-Coder-V2.5-Dev

Kwaipilot (Kuaishou)🇨🇳Major lab

TextCodeopen

This repository contains the model weights and configuration files for the post-trained KAT-Coder-V2.5-Dev in the Hugging Face Transformers format. The artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

Instella-MoE

AMD🇺🇸Major lab

Textopen

Instella-MoE✨: Fully Open State-of-the-Art Mixture-of-Experts Language Model

8 variants16B A3B SFT W4A16 Llmcompressor16B A3B SFT W8A8 Llmcompressor16B A3B Think16B A3B SFT16B A3B Pretrain

Solar-Open2

Upstage🇰🇷Major lab

Textopen

Solar-Open2 is a 250 billion parameter large language model designed for agentic workflows, such as office productivity and coding. It utilizes a Hybrid-Attention Mixture-of-Experts architecture, allowing efficient inference even with long context lengths of up to 1 million tokens.

1 variant250B

Gemini 3.6 Flash

Google DeepMind🇺🇸Frontier

TextReasoningVisionAPI

Gemini 3.6 Flash consumes 17 % fewer output tokens than its predecessor and completes multi-step workflows with fewer reasoning steps and tool calls. It targets coding, knowledge work, and agentic automation at lower cost.

1 variantBatch

Gemini 3.5 Flash Lite

Google DeepMind🇺🇸Frontier

TextReasoningVisionAPI

Fastest and most cost-effective 3.5-class model delivering 350 output tokens per second according to the Artificial Analysis Index, significantly outperforming prior Flash-Lite generations in agentic workflows.

1 variantBatch

Cosmos-H-Dreams

NVIDIA🇺🇸Major lab

Videoopen

Cosmos-H-Dreams is a real-time, action-conditioned generative surgical world model that lets a human operator or a learned surgical-robotics policy act inside a synthesized surgical scene and observe the interactions live.

Motif-3

Motif Technologies🇰🇷Major lab

Textopen

Motif 3 Base is the base pretrained checkpoint of Motif 3 — a large-scale, decoder-only Mixture-of-Experts (MoE) language model with 314 billion total parameters and 13.2 billion parameters activated per token. It is built from the ground up by Motif Technologies following a fully in-house, proprietary design.

3 variantsBaseNVFP4Beta

Week of July 13, 202611 models

Gemini 3.5 Flash Cyber

Google DeepMind🇺🇸Frontier

Text

A lightweight cybersecurity model built on Gemini 3.5 Flash and fine-tuned to find, validate, and patch software vulnerabilities in code more efficiently than the mainline Flash models. It is designed for use by security agents scanning large codebases at scale.

Muse Spark 1.1

Meta🇺🇸Frontier

TextReasoningVisionAPI

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...

Shieldstral-1.0

Mistral AI🇫🇷Frontier

Otheropen

Shieldstral is a compact 3B-parameter, policy-adaptive multimodal safety classifier. Instead of predicting a fixed set of moderation categories, Shieldstral evaluates content against a safety policy expressed in natural language and returns a single continuous safety score.

1 variant3B

LLaDA2.2-flash

Ant Group (inclusionAI)🇨🇳Major lab

Textopen

LLaDA2.2-flash is an agent-oriented diffusion language model in the LLaDA2 series.

PII-Tracer

Perplexity🇺🇸Major lab

Embeddingopen

This repository provides PII-Tracer, the detector introduced in PII-TRACE.

2 variantsMLXGGUF

Intern-S2

Shanghai AI Lab (InternLM)🇨🇳Research

TextVisionopen

Intern-S2 is a multimodal foundation model for scientific intelligence and long-horizon agentic tasks. It unifies vision-language pre-training from raw scientific literature pages with large-scale multi-task reinforcement learning across 20+ domains, plus black-box agentic RL in sandboxed environments. Tool calling, images, video, and time-series inputs are supported.

4 variants397B397B FP8Preview 397BPreview 397B FP8

Wan2.2-Animate-2

Alibaba Qwen🇨🇳Frontier

Imageopen

No description yet — the model card or announcement hasn't been summarised.

3 variants14B Distilled Diffusers14B Diffusers14B

Inkling

Thinking Machines🇺🇸Major lab

TextReasoningVisionopenAPI

Inkling ist ein allgemeines Multimodell von Thinking Machines mit 975 Milliarden Parametern (41 Milliarden aktiv), das Text, Bilder und Audio versteht und Text generiert. Es nutzt eine MoE-Architektur mit kontrollierbarer Reasoning-Effort und ist auf agentic Workflows, Coding und wissenschaftliches Reasoning ausgelegt.

2 variantsFreeNVFP4

Hy-Embodied-VLM-1.0

Tencent🇨🇳Major lab

TextVisionopen

Hy-Embodied-VLM-1.0 Efficient Physical-World Agents Tencent Robotics X × Hy Vision Team × Futian Laboratory

Hy-Embodied-RxBrain-1.0

Tencent🇨🇳Major lab

TextVisionopen

RxBrain Embodied Cognition Foundation Model with Joint Language–Visual Reasoning and Imagination Tencent Robotics X × Futian Laboratory × Tencent Hy Team

Ising-Calibration-1.5

NVIDIA🇺🇸Major lab

TextVisionopen

NVIDIA-Ising-Calibration-1.5-31B-BF16 is a dense multimodal vision-language model built on Gemma 4 31B.

2 variants31B NVFP431B BF16

Week of July 6, 202614 models

Wan-Dancer

Alibaba Qwen🇨🇳Frontier

Videoopen

Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation

1 variant14B

KAT-Coder-Pro V2.5

Kwaipilot (Kuaishou)🇨🇳Major lab

TextCodeAPI

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

⧉ 262K$0.74 / $2.96OR

GPT-5.6 Luna Pro

OpenAI🇺🇸Frontier

TextReasoningVisionAPI

GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoningreasoning-mode

1 variantBatch
Same release

Grok 4.5

xAI🇺🇸Frontier

TextReasoningVisionAPI

Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.

Aion-3.0-Mini

AionLabs🇮🇱Major lab

TextReasoningAPI

Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each...

Same release
⧉ 131K$0.7 / $1.40OR

Cosmos3-Super-Image2Video-4Step

NVIDIA🇺🇸Major lab

Videoopen

NVIDIA Cosmos™ is a world foundation model platform designed to accelerate the development of Physical AI by enabling machines to understand, simulate, and interact with the physical world across robotics, autonomous driving, and smart space environments, including industrial and factory-scale applications.

Cosmos3-Super-Text2Image-4Step

NVIDIA🇺🇸Major lab

Imageopen

NVIDIA Cosmos™ is a world foundation model platform designed to accelerate the development of Physical AI by enabling machines to understand, simulate, and interact with the physical world across robotics, autonomous driving, and smart space environments, including industrial and factory-scale applications.

Nemotron-Labs-Audex

NVIDIA🇺🇸Major lab

Textopen

We're excited to introduce Nemotron-Labs-Audex-2B, a unified audio-text LLM with a similar recipe as Nemotron-Labs-Audex-30B-A3B. Audex-2B extends the vocabulary for discrete audio tokens used for speech and general audio outputs, as well as an audio encoder for speech and general audio inputs.

2 variants2B30B A3B

Week of June 29, 202610 models

LongCat 2.0

Meituan LongCat🇨🇳Major lab

TextReasoningopenAPI

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...

2 variantsINT8FP8

PAR

ByteDance Seed🇨🇳Major lab

Otheropen

Three 3-scale PAR model variants are provided.

Laguna S 2.1

Poolside🇺🇸Major lab

TextReasoningopenAPI

Laguna S 2.1 is the latest coding agent model from Poolside. Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

6 variantsFreeNVFP4 MLXGGUFINT4NVFP4

Leanstral-1.5

Mistral AI🇫🇷Frontier

Otheropen

Leanstral 1.5 is an open-source code agent model designed for Lean 4, a proof assistant capable of expressing complex mathematical objects such as perfectoid spaces and software specifications like properties of Rust fragments.

1 variant119B A6B

Cosmos3-Edge

NVIDIA🇺🇸Major lab

Otheropen

NVIDIA Cosmos3-Edge is a 4B-parameter omni-modal world model that generates text, images, video, and action commands from multimodal inputs. Its Mixture-of-Transformers architecture combines autoregressive decoding for text with diffusion-based denoising for other modalities. Optimized for physical AI, it serves robotics, autonomous vehicles, and smart space simulations.

Claude Sonnet 5

Anthropic🇺🇸Frontier

TextReasoningVisionAPI

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

1 variantBatch

Nano Banana 2 Lite

Google DeepMind🇺🇸Frontier

TextReasoningVisionImageAPI

A text-to-image generation and editing model used for creating graphics, posters, and social-media content. It supports object segmentation, direct in-image text editing, and is integrated into Google Workspace apps.

Nemotron-Parse-2.0

NVIDIA🇺🇸Major lab

TextVisionopen

NVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information.

Moondream3.1

Moondream🇺🇸Major lab

TextVisionopen

Moondream 3.1 is a vision language model with a mixture-of-experts architecture (9B total parameters, 2B active). It delivers state-of-the-art visual reasoning and detection while staying fast and cheap to deploy.

1 variant9B A2B

Tabfm 1.0.0

Google DeepMind🇺🇸Frontier

Otheropen

TabFM is a zero-shot tabular foundation model from Google Research.

2 variantsJaxPytorch

Week of June 22, 20264 models

OlmoEarth-v1_2

Allen Institute for AI (Ai2)🇺🇸Research

Embeddingopen

OlmoEarth-v11-Base is a ViT-Base (114M parameters) foundation model for remote sensing tasks over Sentinel-1, Sentinel-2, and Landsat images or image time series.

1 variantBase

Fugu Ultra

◆ Underrated

Sakana AI🇯🇵Major lab

TextVision

Sakana AI's higher-performance Fugu model — not a monolithic LLM but a learned multi-agent orchestration system that routes requests across multiple models and tools. Multimodal (text+image input), 1M context, $5/$30 per MTok on OpenRouter.

Nemotron-Labs-3-Puzzle

NVIDIA🇺🇸Major lab

Textopen

Nemotron-Labs-3-Puzzle-75B-A9B is a deployment-optimized large language model developed by NVIDIA, derived from Nemotron-3-Super-120B-A12B.

3 variants75B A9B FP875B A9B BF1675B A9B NVFP4

AgentWorld

Alibaba Qwen🇨🇳Frontier

Textopen

Qwen-AgentWorld is the first language world model to cover seven agent interaction domains within a single model. It simulates agentic environments via long chain-of-thought reasoning, predicting the next environment state given an agent's action and interaction history.

1 variant35B A3B

Week of June 15, 20269 models

GLM-5.2

◆ Underrated

Z.ai (Zhipu)🇨🇳Frontier

TextReasoningopenAPI

GLM-5.2 is an open-source model from Z.AI for long-horizon tasks, with a stable 1-million-token context and strong agentic coding. It combines code generation with tool use and supports adjustable thinking-effort levels. It scores 62.1% on SWE-bench Pro, sits close to the frontier on AIME and GPQA-Diamond, and improves markedly over GLM-5.1. The MoE architecture with IndexShare cuts compute per token by a factor of 2.9.

1 variantFP8

Laguna 2.1

Poolside🇺🇸Major lab

TextReasoningopenAPI

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

12 variantsXSXS FreeXS NVFP4 MLXXS GGUFXS Dflash NVFP4

Unlimited-OCR

Baidu🇨🇳Major lab

TextVisionopen

Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.3 + CUDA12.9:

Nano Banana Pro

Google DeepMind🇺🇸Frontier

TextReasoningVisionImageAPI

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

Transcribe Arabic

Cohere🇨🇦Major lab

Audioopen

No description yet — the model card or announcement hasn't been summarised.

1 variant202607

Gemini 3.1 Flash Image

Google DeepMind🇺🇸Frontier

VisionImageAPI

Google's latest Gemini Flash generation with native image input/output. Available on OpenRouter ($0.50/$3 per MTok). Gemini 3.x is a standalone series alongside Gemini 2.5.

FastContext 1.0

◆ Underrated

Microsoft🇺🇸Major lab

Codeopen

Microsoft's specialized code-search model for coding agents — powers the 'Explore' subagent in SWE-FastContext. Not a general-purpose LLM, but relevant for agentic setups.

1 variant4B

North Mini Code 1.0

◆ Underrated

Cohere🇨🇦Major lab

CodeopenAPI

North Mini Code is an open-weight sparse MoE model (30B total, 3B active) specialized for code generation, agentic software engineering, and terminal tasks. It supports tool-use through chat templates and was post-trained with supervised fine-tuning followed by reinforcement learning with verifiable rewards focused on agentic coding.

2 variantsW4A16FP8

MolmoMotion-H3-F30

Allen Institute for AI (Ai2)🇺🇸Research

TextVisionopen

MolmoMotion is a 4B vision-language model that forecasts 3D point trajectories under natural-language action instructions.

1 variant4B

Week of June 8, 202614 models

Kimi K3

Moonshot AI🇨🇳Frontier

TextReasoningVisionopenAPI

Kimi K3 ist ein Open-Weight-Multimodalmodell von Moonshot AI mit 2,8 Billionen Parametern (davon 104 Mrd. aktiv) im MoE-Design und nativer Bild- sowie Videoerkennung. Die auf Kimi Delta Attention und Attention Residuals aufbauende Architektur zielt auf langfristiges Coding, Reasoning und agentic Knowledge Work ab; das Modell unterstuetzt einen Kontextfenster von 1 Million Tokens und versteht Text, Bilder und Videos im selben Modell.

1 variantBatch

Kimi K2.7 Code

Moonshot AI🇨🇳Frontier

ReasoningCodeopenAPI

Kimi K2.7 Code is Moonshot AI's coding-specialised agentic model. It accepts image input and tool calls and can handle complex software-engineering workflows end to end. It scores 62.0 on Kimi Code Bench v2 and uses about 30% fewer thinking tokens than its predecessor K2.6. It trails GPT-5.5 slightly on coding benchmarks but beats Claude Opus 4.8 on MCP Mark Verified with 81.1.

VISTA

Ant Group (inclusionAI)🇨🇳Major lab

TextVisionopen

VISTA-9B are GUI-grounding vision-language models trained from Qwen3.5 9B backbones with VISTA: View-Consistent Self-Verified Training for GUI Grounding.

2 variants9B4B

LFM2.5-Encoder

Liquid AI🇺🇸Major lab

TextEmbeddingopen

LFM2.5-Encoder is a family of multilingual bidirectional encoders built on the LFM2 architecture, available in two sizes:

7 variants230M350M350M Prompt Router350M Policy Linter350M Diffusion

MobileMoE-S

Meta🇺🇸Frontier

Textopen

No description yet — the model card or announcement hasn't been summarised.

2 variantsBaseSFT

ZONOS2

Zyphra🇺🇸Major lab

Audioopen

ZONOS2 is our latest text-to-speech model trained on more than 6 million hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers at low latency with MoE. ZONOS2 excels at high-fidelity and naturalistic voice cloning.

1 variantGGUF

EvoQuality

ByteDance Seed🇨🇳Major lab

TextVisionopen

- Model Name: EvoQuality (Self-Evolving VLM for Image Quality Assessment) - Task: No-Reference Image Quality Assessment (NR-IQA), supporting both single-image quality scoring and pairwise quality comparison (ranking) - Core Idea: Without relying on any human-annotated quality scores or distortion-type labels,…

UltraX

OpenBMB🇨🇳Research

Textopen

UltraX is a function-calling refinement framework for large-scale pre-training data.

1 variant0.6B Preview

Claude Fable 5

Anthropic🇺🇸Frontier

ReasoningVisionAPI

Anthropic's new top flagship — replaces the Opus tier as the strongest generally available model. Arena Elo 1508 (rank 1). 1M context, Adaptive Thinking always on, $10/$50 per MTok. Fable 5 + Mythos 5 (invite-only) launched together on June 9, 2026.

Diffusiongemma

Google DeepMind🇺🇸Frontier

TextVisionopen

DiffusionGemma is a generative model built by Google DeepMind. Based on the 26B A4B Mixture-of-Experts (MoE) Gemma 4 architecture, DiffusionGemma generates tokens using discrete diffusion.

1 variant26B A4B IT

SCAIL-2

Z.ai (Zhipu)🇨🇳Frontier

Videoopen

SCAIL-2 is an open-source model for end-to-end controlled character animation. It animates a reference character with a driving video, and also supports character replacement and multi-character scenarios without relying on intermediate pose representations.

PP-OCRv6_tiny

Baidu🇨🇳Major lab

TextVisionopen

PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks

4 variantsRec OnnxDet OnnxDet SafetensorsRec Safetensors

PP-OCRv6_small

Baidu🇨🇳Major lab

TextVisionopen

PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks

4 variantsRec OnnxDet OnnxDet SafetensorsRec Safetensors
Same release

Week of June 1, 202611 models

Nemotron 3 Ultra

NVIDIA🇺🇸Major lab

ReasoningTextopenAPI

NVIDIA's powerful open-weights MoE model: hybrid LatentMoE + MTP layers, 1M context, 550B/55B active. Free on OpenRouter. Focus: complex multi-agent workflows, code, math, science.

3 variants550B550B A55B NVFP4550B A55B Base BF16

Nemotron 3.5 Content

NVIDIA🇺🇸Major lab

TextReasoningVisionAPI

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

2 variantsSafetySafety Free

ArtiFixer

NVIDIA🇺🇸Major lab

Otheropen

ArtiFixer is a few-step causal auto-regressive model that enhances and extends 3D reconstruction. The related source code provides implementations for training, evaluation, and inference, supporting various stages including bidirectional training, diffusion forcing, and Self-Forcing-style DMD distillation.

Nex-N2-Pro

◆ Underrated

Nex AGI (Shanghai Innovation Inst.)🇨🇳

ReasoningTextopenAPI

397B model from the Shanghai Innovation Institute — open source, Apache 2.0, free on OpenRouter. Agentic-focused, with papers on a 'Unified Ecosystem for Large-Scale Environment Construction'. 7.87K HF likes despite barely any Western coverage.

Qwen3.7 Plus

Alibaba Qwen🇨🇳Frontier

ReasoningTextAPI

Qwen's latest Plus tier — 1M context, $0.32/$1.28 per MTok. Qwen3.7 Max (stronger) and Plus (more efficient) form the top of the Qwen3.7 line.

LFM2.5

Liquid AI🇺🇸Major lab

TextReasoningopenAPI

LFM2.5 is a hybrid model designed for on-device deployment, featuring 230 million parameters and optimized for fast edge inference. It is suitable for general-purpose text generation and agentic tasks, but not recommended for reasoning-heavy workloads.

34 variants2.6B Free8B A1B Dspark GGUF2.6B Dspark GGUF1.2B Instruct Dspark GGUF2.6B Dspark

MiniMax M3

◆ Underrated

MiniMax🇨🇳Major lab

VisionReasoningopenAPI

MiniMax is barely known in the West — but its M3 model on OpenRouter ($0.30/$1.20 per MTok) is one of the cheapest 1M-context multimodal providers. Successor to Text-01 (4M context).

1 variantMXFP8

Moshika

Kyutai🇫🇷Research

Audioopen

No description yet — the model card or announcement hasn't been summarised.

1 variantRL Seamless

DeepSeek V4 Pro

DeepSeek🇨🇳Frontier

ReasoningTextopenAPI

DeepSeek's heavyweight: 1.6T total, 49B active, 1M context, FP4/FP8 mixed precision. Codeforces 3206 Elo beats GPT-5.4. MIT license, fully open weights. $0.435/$0.87 per MTok on OpenRouter.

3 variants081308130813 Batch
Same release

Bernini-R

ByteDance Seed🇨🇳Major lab

Otheropen

- [2026-06-09] We open-sourced the 1.3B weights of the Bernini Renderer (Bernini-R) on ByteDance/Bernini-R-1.3B-Diffusers.

2 variants1.3B DiffusersDiffusers

Week of May 25, 20262 models

Claude Opus 4.8

Anthropic🇺🇸Frontier

ReasoningVisionAPI

Current Opus-tier model (until Claude Fable 5). Arena Elo ~1490. 1M context, Adaptive Thinking, knowledge through Jan 2026. $5/$25 per MTok direct, similar on OpenRouter.

Step 3.7 Flash

◆ Underrated

StepFun🇨🇳Major lab

VisionReasoningopenAPI

201B multimodal open-weights model from StepFun — underrated in the West. Image+text input, strong reasoning. Successor to Step 3.5 Flash (199B). $0.20/$1.15 per MTok on OpenRouter.

Week of May 18, 20264 models

Qwen3.7 Max

Alibaba Qwen🇨🇳Frontier

ReasoningTextAPI

Strongest Qwen3.7 variant: 1M context, $1.25/$3.75 per MTok. More capable than Plus, but pricier. The Qwen3.7 line shipped late May 2026 as the successor to Qwen3.6.

Grok Build 0.1

◆ Underrated

xAI🇺🇸Frontier

CodeReasoningAPI

xAI's coding-focused model — 256K context, $1/$2 per MTok. 'Build' implies software development as the main use case. A separate coding line alongside the general-purpose Grok 4.3.

Gemini 3.5 Flash

Google DeepMind🇺🇸Frontier

ReasoningVisionAPI

Google's current Flash generation: 1M context, all modalities, $1.50/$9 per MTok on OpenRouter. Gemini 3.5 Flash is the speed tier of the Gemini 3.x series that replaces Gemini 2.5.

Kimi K2.6

Moonshot AI🇨🇳Frontier

ReasoningVisionopenAPI

Multimodal predecessor of K2.7-Code — Visual Agentic Intelligence with 1.1T parameters. 2.66M HF downloads. Still relevant for multimodal tasks (images, video) since K2.7 is code-only.

Week of May 11, 20261 model

GLM-5.1

◆ Underrated

Z.ai (Zhipu)🇨🇳Frontier

TextReasoningopen

Direct predecessor of GLM-5.2 (June 20). Still interesting for local deployment. The GLM-5.1 FP8 variant has 1.33M downloads — massive interest from the CN community.

Week of May 4, 20261 model

Gemini 3.1 Flash Lite

Google DeepMind🇺🇸Frontier

VisionTextAPI

Google's cheapest Gemini 3.x option: 1M context, $0.25/$1.50 per MTok. The 'Lite' variant is the first choice when token costs are critical and multimodality is needed.

Week of April 27, 202611 models

Laguna M.1

◆ Underrated

Poolside🇺🇸Major lab

CodeReasoningopenAPI

Poolside's strong coding-agent model: 226B MoE, Apache 2.0, SWE-bench Pro 49.2%. Free on OpenRouter! Specialized for agentic long-horizon coding tasks. Poolside is a little-known US AI company.

1 variantBase

Laguna XS.2

◆ Underrated

Poolside🇺🇸Major lab

CodeopenAPI

Poolside's lean coding model: 33B MoE, 3B active, for local deployments. Free on OpenRouter. Sliding-window attention for very fast inference. 231K downloads on HF.

1 variantXS

Grok 4.3

xAI🇺🇸Frontier

ReasoningVisionAPI

xAI's current frontier model on OpenRouter ($1.25/$2.50 per MTok). Grok 4.3 positions itself as a strong all-rounder. Also accessible via Grok.com / X Premium.

Mistral Small 4

Mistral AI🇫🇷Frontier

TextReasoningopenAPI

Mistral ironically calls 119B 'Small' — Apache 2.0, with instruction-following, reasoning and coding in one. Mistral positions itself as the strongest European open-weights challenger.

1 variant119B

Mistral Medium 3.5

Mistral AI🇫🇷Frontier

TextVisionAPI

Mistral Medium 3.5 is Mistral's first flagship to handle instruction-following, reasoning and coding in a unified way. 418K HF downloads, EAGLE-acceleration variant available. $1.50/$7.50 per MTok on OpenRouter.

Granite 4.1

◆ Underrated

IBM🇺🇸Major lab

TextCodeopenAPI

IBM's Granite 4.x generation: 8B, Apache 2.0, $0.05/$0.10 per MTok (one of the cheapest). Enterprise-focused, strong on structured tasks. IBM ships Granite models consistently with open weights.

1 variant8B

Kimi K2.5

Moonshot AI🇨🇳Frontier

ReasoningVisionopen

Moonshot's 'Visual Agentic Intelligence' — 1.81M HF downloads. Basis for K2.6 and K2.7. One of the first large models with real multimodal agentics at 1T parameters.

Command A+

◆ Underrated

Cohere🇨🇦Major lab

ReasoningVisionopenAPI

Cohere's first multimodal open-weights flagship: 218B MoE, 25B active, 128K context, 48 languages, Apache 2.0. Reasoning tokens, tool use with JSON schema. Not on OpenRouter yet.

Nemotron 3 Nano Omni

◆ Underrated

NVIDIA🇺🇸Major lab

ReasoningVisionopenAPI

NVIDIA's small multimodal reasoning model: 30B MoE, only 3B active, text+image+audio input. Free on OpenRouter. Designed for on-device, 4x faster than its predecessor, reasoning ON/OFF mode.

1 variant30B

Qwen3.6 Flash

Alibaba Qwen🇨🇳Frontier

ReasoningTextAPI

Fast variant of the Qwen3.6 family: 1M context, $0.19/$1.13 per MTok. Qwen3.6 shipped as Flash, 27B, 35B and Max-preview variants — all on April 28, 2026.

Same release

Week of April 20, 20268 models

GPT-5.5

OpenAI🇺🇸Frontier

ReasoningVisionAPI

OpenAI's current frontier model: 'A new class of intelligence for coding and professional work.' 1M context, knowledge cutoff Dec 2025, $5/$30 per MTok. GPT-5.5 Pro as the stronger variant ($30/$180 per MTok).

Same release

MiMo V2.5 Pro

◆ Underrated

Xiaomi MiMo🇨🇳Major lab

ReasoningTextopenAPI

Xiaomi's reasoning model on OpenRouter ($0.435/$0.87 per MTok). MiMo is Xiaomi's first serious LLM push — barely noticed in the West. 1M context. V2.5 (lighter) and V2.5-Pro available.

1 variantFP4 Dflash
Same release

Hy3

◆ Underrated

Tencent🇨🇳Major lab

TextReasoningopenAPI

Tencent's Hunyuan 3 in preview on OpenRouter ($0.063/$0.21 per MTok). One of the cheapest models available. Tencent is a heavyweight that gets little Western attention in AI.

2 variantsPreviewFP8

Ling-2.6

◆ Underrated

Ant Group (inclusionAI)🇨🇳Major lab

ReasoningTextopenAPI

Ant Group's (Alipay's parent) 1-trillion-parameter MoE model — Apache 2.0. Extremely cheap on OpenRouter ($0.075/$0.625 per MTok). 472 likes on HF despite barely any Western coverage.

2 variants1T1T Base
Same release

GPT-5.4 Image 2

OpenAI🇺🇸Frontier

ImageVisionAPI

OpenAI's second image-generation iteration on the GPT-5.4 base: $8/$15 per MTok (input/output). The current standard route for professional AI image generation via OpenRouter API.

Week of March 30, 20261 model

GLM-5

◆ Underrated

Z.ai (Zhipu)🇨🇳Frontier

TextReasoningopen

First release of the GLM-5 generation (April 2026). 2.1K likes on HuggingFace. Zhipu's development rhythm is remarkable: one major release per month.

Week of February 23, 20261 model

Gemma 3n

◆ Underrated

Google DeepMind🇺🇸Frontier

Visionopen

Google's on-device multimodal model: handles text, image, video AND audio at just 4B effective parameters. MatFormer architecture enables sub-models within the same checkpoint. Designed for edge deployments.

1 variantE4B

Week of January 26, 20261 model

Claude Opus 4.7

Anthropic🇺🇸Frontier

ReasoningVisionAPI

A few generations before Fable 5 — but in thinking mode Arena Elo 1502, rank 3 worldwide. Still relevant for setups where Extended Thinking is desired and Fable 5's Adaptive Thinking isn't enough.

Week of October 27, 20251 model

GPT-5.4

OpenAI🇺🇸Frontier

ReasoningVisionAPI

Predecessor of GPT-5.5, but still very widely used ($2.50/$15 per MTok vs. $5/$30 for 5.5). Mini variant (400K context, $0.75/$4.50) for coding agents and subagents. IMOAnswerBench 91.4%.

Week of April 28, 20251 model

Qwen3

◆ Underrated

Alibaba Qwen🇨🇳Frontier

ReasoningTextopen

Alibaba's strong open-weights model from April 2025 — basis for many downstream variants. Qwen3.5 and 3.6 followed in 2026. Often overlooked in the West despite AIME 2025 85.7%.

1 variant235B A22B

Week of March 31, 20251 model

Llama 4 Maverick

Meta🇺🇸Frontier

Visionopen

Meta's multimodal MoE flagship: 17B active, 128 experts, 402B total, 1M-token context. Scout variant (16E, 109B) for lean deployments. Llama 5 is expected in H2 2026.

Week of March 10, 20251 model

Gemma 3

Google DeepMind🇺🇸Frontier

Visionopen

Google's open-weights flagship: 27B (also 1B, 4B, 12B), 128K context, text+image, 140+ languages. Gemma 3 is the basis for many community fine-tunes. Superseded by Gemma 4.

1 variant27B