Releases from frontier labs, major companies and research institutes — grouped by model, deduplicated across HuggingFace, OpenRouter and official blogs.
Releases per weekFrontierMajor labResearch
Click a week to filter
Tier
Type
Source209 models
Last week20 models
XM
MiMo-V2.6-Flash-MOPD
Xiaomi MiMo·🇨🇳·Major lab
Textopen
MiMo-V2.6-Flash-MOPD is an open-weight multimodal model by Xiaomi supporting text, images, audio, and video with tool-use capability and structured reasoning via <think> tags. Built on a 309B-parameter sparse MoE with 15B activated tokens per step and 1M token context, it addresses tool-call repetition from its predecessor through MOPD distillation using domain-specialized teachers.
Same release
GD
Gemini 3.8 Live
Google DeepMind·🇺🇸·Frontier
Text
Gemini 3.8 Live erweitert ein Live-Dialytemodell um eine Videoavatar-Funktion, die Sprache und visuelles Feedback simultan verarbeitet — mit Lippen-Sync, multi-modalen Eingaben (Video plus Audio) und asynchronem Tool-Calling waehrend des Dialoges. Es laeuft nahtlos ueber 97 Sprachen hinweg ohne Abbruch der Videofidelitaet.
ZA
GLM 5.3 Prime
Z.ai (Zhipu)·🇨🇳·Frontier
TextReasoningAPI
GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. It supports text input and output with a 1M-token...
AQ
Qwen3.8 Max Prime
Alibaba Qwen·🇨🇳·Frontier
TextReasoningVisionAPI
Qwen3.8 is a causal language model with vision capabilities, designed for complex tasks involving coding, professional work, and research. It supports a context length of up to 1,000,000 tokens and is compatible with various deployment tools.
AI
Aion 3.5 Mini
AionLabs·🇮🇱·Major lab
TextReasoningAPI
Aion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, lower-cost sibling of Aion 3.5 and uses...
Same release
UP
Solar Mini 4
Upstage·🇰🇷·Major lab
TextReasoningAPI
Solar 4 is a model family by Upstage. The Pro variant supports a context window of 512K tokens and is available via OpenRouter under the ID upstage/solar-pro4.
NV
NV-Reason-CT Open 3D CT VLM
NVIDIA·🇺🇸·Major lab
TextVisionopen
NV-Reason-CT processes 3D CT volumes natively (not slice-by-slice) and generates structured diagnostic reports with step-by-step chain-of-thought reasoning validated by NIH radiologists. It scores state-of-the-art on CT-RATE (Macro-F1 0.614).
?
Space Bunny Alpha
◌ Stealth
Stealth·🌐·Stealth
TextReasoningVisionAPI
Space Bunny Alpha is an anonymous large model with blazing-fast inference, strong coding capabilities and native multimodal input support. It delivers adjustable reasoning effort, and a 1M-token context window.
1 variantAlpha
OP
GPT-6 Luna Pro
OpenAI·🇺🇸·Frontier
TextReasoningVisionAPI
GPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoningreasoning-mode
1 variantBatch
Same release
AN
Claude Opus 5.5
Anthropic·🇺🇸·Frontier
TextReasoningVisionAPI
Anthropic's flagship model with significantly improved price–performance ratio: it communicates more directly than its predecessor, matches Fable 5.1's intelligence at lower cost, and achieves higher token efficiency, making it suitable for extended reasoning and agentic workflows.
1 variantBatch
XA
Grok 4.7
xAI·🇺🇸·Frontier
TextReasoningVisionAPI
Anthropic's Opus-tier flagship addresses prior feedback by communicating clearly; delivers Fable 5.1 intelligence at a 20% lower price and is notably token-efficient, especially for long agentic workflows via slashed cache prices.
AQ
Qwen3.8 Omni Flash
Alibaba Qwen·🇨🇳·Frontier
TextReasoningVisionAPI
Qwen3.8 is a causal language model with vision capabilities, designed for complex tasks involving coding, professional work, and research. It supports a context length of up to 1,000,000 tokens and is compatible with various deployment tools.
XM
MiMo-V2.6-Pro-UltraSpeed
Xiaomi MiMo·🇨🇳·Major lab
TextReasoningVisionAPI
MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro checkpoint, it matches the original model in quality while delivering roughly 10x...
XM
MiMo-V2.6-Flash
Xiaomi MiMo·🇨🇳·Major lab
TextReasoningVisionopenAPI
A model variant within the MiMo-V2.6 family using a sparse Mixture-of-Experts architecture with 15 billion active parameters out of 309 billion total. Processes text, images, video, and audio in one model. Trained with groupwise reinforcement learning designed to expand capabilities through self-improvement — targeting complex agentic tasks, tool use, and multi-session long contexts.
1 variantRL
Same release
Week of September 14, 20265 models
ZA
GLM 5.3 FlashX
Z.ai (Zhipu)·🇨🇳·Frontier
TextReasoningVisionAPI
GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. Built on the same hybrid sparse and linear attention architecture...
AG
Ming-Image-0.1-Design
Ant Group (inclusionAI)·🇨🇳·Major lab
Imageopen
Ming-Image-0.1-Design is a 6B text-to-image model for UI, infographics, posters, and other text-rich visual designs. It generates complete visual compositions and supports RGBA output with transparent backgrounds.
AG
Ming-Image-0.1-Design-Layer
Ant Group (inclusionAI)·🇨🇳·Major lab
Otheropen
Ming-Image-0.1-Design-Layer decomposes a flattened design image into a requested number of RGBA layers using an image and a layer plan.
AG
Realtime-Venus
Ant Group (inclusionAI)·🇨🇳·Major lab
TextVisionopen
A full-duplex interaction system with asynchronous delegation
AQ
Qwen-Image-2.1
Alibaba Qwen·🇨🇳·Frontier
Imageopen
Qwen-Image-2.1 is a 7B-parameter image-generation and editing model supporting native RGBA output, multi-reference edits (up to 10), and resolutions from 1536x2752 up to 2752x1536.
2 variantsPE I2iPE T2i
Week of September 7, 20266 models
SA
Fugu Ultra v2
Sakana AI·🇯🇵·Major lab
TextReasoningVisionAPI
Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to...
SA
Fugu Max
Sakana AI·🇯🇵·Major lab
TextReasoningVisionAPI
Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent orchestration system: a language model trained to route...
SA
Atria-Dawn
Shanghai AI Lab (InternLM)·🇨🇳·Research
Otheropen
Atria Dawn Preview is a preview release of a new-generation agentic model developed by the Shanghai Artificial Intelligence Laboratory.
3 variantsPreview Ascend W8A8Preview FP8Preview
DE
DeepSeek V4.1 Flash
DeepSeek·🇨🇳·Frontier
TextReasoningVisionopenAPI
Multimodal MoE model that natively processes images and text with a 1M-token context window. Based on a 552B backbone, it activates 8B/16B per step, cutting KV-cache overhead to roughly a quarter of its predecessor.
1 variantBatch
OP
ChatGPT Images 2.5
OpenAI·🇺🇸·Frontier
Image
ChatGPT Images 2.5 helps turn your ideas, sketches, and reference photos into more personalized, polished images that better reflect your ideas.
IN
Mercury 2.5
Inception·🇺🇸·Major lab
TextReasoningAPI
Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
Week of August 31, 202616 models
AG
LLaDA-UI
Ant Group (inclusionAI)·🇨🇳·Major lab
TextVisionopen
LLaDA-UI is an MoE-based, block-wise diffusion vision-language GUI agent. It understands screenshots at their native aspect ratio and produces grounded coordinates or structured actions for mobile, desktop, and web interfaces.
AG
LLaDA2.2-mini
Ant Group (inclusionAI)·🇨🇳·Major lab
Textopen
LLaDA2.2-mini is the lightweight variant of the agentic diffusion language model in the LLaDA2 series.
OP
GPT-6 Astra
OpenAI·🇺🇸·Frontier
TextReasoningVisionAPI
GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
1 variantBatch
Same release
AG
Ling 3.0 Flash VL
Ant Group (inclusionAI)·🇨🇳·Major lab
TextReasoningVisionopenAPI
Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while adding native visual perception and advanced visual...
3 variantsINT4FP4FP8
AG
Ling 3.0 Flash Sante
Ant Group (inclusionAI)·🇨🇳·Major lab
TextReasoningAPI
Ling 3.0 Flash Sante is a health and medicine-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for...
1 variantFree
TE
EVIE
Tencent·🇨🇳·Major lab
Otheropen
EVIE-8B ist ein vision-language Modell für visuelle Dokumentensuche und -abfrage basierend auf Qwen3.5 mit bidirektionaler Selbstachtund Multi-Vector-Embeddings. Es dient als Distillierungslehrer für die kompakte EVIE-4.5B-Variante.
2 variants8B4.5B
AQ
Qwen3.8 Max
Alibaba Qwen·🇨🇳·Frontier
TextReasoningVisionAPI
Qwen3.8 is a causal language model with vision capabilities, designed for complex tasks involving coding, professional work, and research. It supports a context length of up to 1,000,000 tokens and is compatible with various deployment tools.
GD
WeatherNext 3
Google DeepMind·🇺🇸·Frontier
Other
AI weather model that uses real-time satellite data instead of traditional physics simulations to produce hourly forecasts at five times the resolution of its predecessor. Also tracks precipitation and includes clean-energy variables.
ME
Muse Spark 1.3
Meta·🇺🇸·Frontier
TextReasoningVisionAPI
Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...
1 variantContributor
GD
Gemini 3.8 Flash
Google DeepMind·🇺🇸·Frontier
TextReasoningVisionAPI
Gemini 3.8 Live adds near real-time streaming video with expressive avatars — it processes audio and visual inputs simultaneously, supports precise lip-syncing across 97 languages, and can execute tool calls asynchronously while maintaining an uninterrupted conversation.
1 variantBatch
MI
VibeVoice-ASR-Streaming
Microsoft·🇺🇸·Major lab
Audioopen
VibeVoice-ASR-Streaming is a unified streaming ASR model that transcribes Who (Speaker) said What (Content), with support for Customized Hotwords and 10 languages.
2 variants1.5B7B
CO
Tiny Aya En Thinker
Cohere·🇨🇦·Major lab
Textopen
No description yet — the model card or announcement hasn't been summarised.
CO
Tiny Aya L2 Thinker
Cohere·🇨🇦·Major lab
Textopen
No description yet — the model card or announcement hasn't been summarised.
AN
Claude Fable 5.1
Anthropic·🇺🇸·Frontier
TextReasoningVisionAPI
Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
1 variantBatch
JA
Jina OCR V1
Jina AI·🇩🇪·Major lab
TextVisionopen
jina-ocr-v1 is an end-to-end document parsing model designed for high-quality OCR at an efficient serving point.
Week of August 24, 20269 models
AG
LLaDA-Image-Turbo
Ant Group (inclusionAI)·🇨🇳·Major lab
Imageopen
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
1 variantFP8
Same release
AQ
Drive-1.0
Alibaba Qwen·🇨🇳·Frontier
TextVisionopen
Qwen-Drive-1.0 retains the architecture of the pretrained Qwen3.5 vision-language model and integrates 3D perception, visual question answering, and motion planning within a unified framework.
1 variant4B
TE
Hy4
Tencent·🇨🇳·Major lab
TextReasoningopenAPI
Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that...
3 variantsPreviewPreview FP8Preview
AG
Ling 3.0 Flash Fin
Ant Group (inclusionAI)·🇨🇳·Major lab
TextReasoningopenAPI
Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B total. It is designed for real-world investment...
4 variantsFreeFP4INT4FP8
OP
MiniCPM5
OpenBMB·🇨🇳·Research
Textopen
MiniCPM5-2B ist ein dichter 2B-Transformer von OpenBMB, optimiert für On-Device-Einsatz und ressourcenbeschränkte Szenarien. Das Modell konkurrenzieren 4B-Klassenmodelle und zeigt Vorteile in Code, Mathematik, Langkontext-Verständnis sowie Tool-Use und agentic Tasks.
9 variants2B2B Dspark GGUF2B GPTQ2B Dspark2B GGUF
ZA
GLM 5.3 Flash
Z.ai (Zhipu)·🇨🇳·Frontier
TextReasoningVisionopenAPI
GLM 5.3 is a language model developed by Z.AI (Zhipu) with a large context window of 1,048,576 tokens.
2 variantsBatchBF16
AQ
Qwen3.8 Flash
Alibaba Qwen·🇨🇳·Frontier
TextReasoningVisionopenAPI
Qwen3.8 is a causal language model with vision capabilities, designed for complex tasks involving coding, professional work, and research. It supports a context length of up to 1,000,000 tokens and is compatible with various deployment tools.
1 variantFP8
NV
Nemotron-3-Diarization
NVIDIA·🇺🇸·Major lab
Otheropen
Nemotron 3 Diarization is an open-weight speaker diarization model designed to determine "who spoke when" in real-world audio. It supports both streaming and offline inference and handles up to eight speakers.
1 variantPreview
Week of August 17, 20264 models
DE
DeepSeek V4 Flash Vision
DeepSeek·🇨🇳·Frontier
TextReasoningVisionopenAPI
Multimodal text-and-image model with substantially improved agent capabilities versus its text-only predecessor; competes in text-only agent tasks but trails leading proprietary models.
2 variantsExpExp
TE
AuK-Flash
Tencent·🇨🇳·Major lab
Audioopen
- [2026/09/09] 🎉 AuK is now open-source. Code and model weights are publicly available.
Same release
ZA
GLM 5.3
Z.ai (Zhipu)·🇨🇳·Frontier
TextReasoningopenAPI
GLM 5.3 is a language model developed by Z.AI (Zhipu) with a large context window of 1,048,576 tokens.
2 variantsBatchBF16
Week of August 10, 202611 models
CO
North-Small-Translate-1.0
Cohere·🇨🇦·Major lab
Textopen
No description yet — the model card or announcement hasn't been summarised.
2 variantsW4A16FP8
GD
Gemini 3.7 Flash
Google DeepMind·🇺🇸·Frontier
TextReasoningVisionAPI
Gemini 3.7 is a model developed by Google DeepMind, featuring two variants: Flash and Flash Batch. It has a context window of 1,048,576 tokens.
1 variantBatch
NV
Cmd
NVIDIA·🇺🇸·Major lab
Videoopen
Hmrishav Bandyopadhyay 1,2 , Xuanchi Ren 1 , Zijian Huang 1 , Jay Zhangjie Wu 1 , Tianshi Cao 1 , Ruilong Li 1 , Bryan Chu 1 , Sanja Fidler 1 , Yi-Zhe Song 2 , Zian Wang 1
XA
Grok 4.6
xAI·🇺🇸·Frontier
TextReasoningVisionAPI
Grok 4.6 is a model developed by xAI, featuring a context window of 500,000 tokens. It is available in a standard variant.
BS
Seed 2.1 Turbo
ByteDance Seed·🇨🇳·Major lab
TextReasoningVisionAPI
Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software delivery, multi-step task execution, and understanding visual and...
BS
Seed-2.0-Code
ByteDance Seed·🇨🇳·Major lab
TextReasoningVisionCodeAPI
Seed-2.0 is a model developed by ByteDance Seed, designed for code-related tasks with a large context window.
SA
Namazu
Sakana AI·🇯🇵·Major lab
TextReasoningVisionAPI
Namazu is a model from Sakana AI with a 256K token context window (262144). No information is available on parameter count, licensing, weight openness, intended use cases, or capability tier.
UP
Solar Pro 4
Upstage·🇰🇷·Major lab
TextReasoningAPI
Solar 4 is a model family by Upstage. The Pro variant supports a context window of 512K tokens and is available via OpenRouter under the ID upstage/solar-pro4.
CO
North-Micro-Vision
Cohere·🇨🇦·Major lab
TextVisionopen
North-Micro-Vision is a 2.4B open-weight vision-language model combining a 400M native-resolution vision encoder with a 2B Command A+ language backbone, designed for prototyping and fine-tuning.
1 variantInstruct
AG
Ling-3.0-tiny
Ant Group (inclusionAI)·🇨🇳·Major lab
Textopen
Ling-3.0-tiny is a 7.9B hybrid MoE language model that activates only 1.3B parameters per token. It combines a 3:1 alternating layout of Kimi Delta Attention and Multi-Head Latent Attention layers with a sparse MoE feed-forward network of 128 experts (8 routed plus 1 shared per token). The model provides reasoning via a configurable thinking mode, supports multi-turn tool calling through XML-tagged function calls, and offers BF16, FP8, and INT4 quantized weights. Benchmarks report around 160+ tokens/s inference speed on an H20 GPU, approximately 8.3 GiB peak VRAM at 8K context in FP8, and an Artificial Analysis Intelligence Index score of 25 and Agentic Index score of 16. SGLang integration includes a built-in speculative decoding recipe (NEXTN/MTP) and YaRoN-based context extension up to 262K tokens. No explicit evaluation table values are published in the released material beyond those two index scores.
7 variantsSingprobeGGUFBaseBase 30TBase Midtrain
LI
LTX-2.5-Pre-Trained
Lightricks·🇮🇱·Major lab
Videoopen
No description yet — the model card or announcement hasn't been summarised.
Week of August 3, 20267 models
ME
Muse Glimmer
Meta·🇺🇸·Frontier
TextReasoningVisionopenAPI
Muse Glimmer is a 30-billion-parameter causal language model designed for local agentic AI workflows, optimized for NVIDIA hardware. It integrates multi-step reasoning, reliable tool use, and multimodal understanding, enabling complex task completion without cloud dependency.
Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows that need multi-step reasoning.
MiniMax Music 3 is a music generation model capable of creating complete songs up to five minutes long from lyrics and a style description. It accepts lyrics with section tags such as verse, chorus, bridge, and outro, along with structured metadata covering genre, tempo, instrumentation, and vocal characteristics. The model produces stereo audio at 32 kHz and maintains thematic coherence across full song structures including intro, development, and finale.
LA
LFM2.5-VL
Liquid AI·🇺🇸·Major lab
TextVisionopen
LFM2.5-VL-3B is a multimodal variant of LFM2.5, a family of hybrid models designed for on-device deployment. It builds on LFM2-VL-3B with further mid- and post-training.
Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...
1 variantContributor
AQ
Qwen3.8
Alibaba Qwen·🇨🇳·Frontier
TextReasoningVisionopenAPI
Qwen3.8 is a causal language model with vision capabilities, designed for complex tasks involving coding, professional work, and research. It supports a context length of up to 1,000,000 tokens and is compatible with various deployment tools.
7 variants27B27B Free2.4T A95B27B27B FP8
IB
Granite Speech 5.0
IBM·🇺🇸·Major lab
Audioopen
Model Summary: Granite Speech 5.0 TurboCTC Non-commercial is a compact 470 million parameter English ASR model with very high inference speed that is released for research and noncommercial purposes only. We invite users to refer to granite-speech-5.0-470m-turboctc for other use cases.
2 variants470M Turboctc Nc470M Turboctc
Week of July 27, 20269 models
NV
Nemotron 3.5 Lightning
NVIDIA·🇺🇸·Major lab
TextReasoningopenAPI
Nemotron 3.5 Lightning ist ein Textmodell von NVIDIA mit einem Kontextfenster von 262 K Tokens. Es ist seit dem 11. August 2026 auf OpenRouter verfuegbar.
LongCat-Flash-Lite-Sparse is a non-thinking Mixture-of-Experts (MoE) model with 69B total parameters and approximately 3B activated parameters per token. Built on LongCat-Flash-Lite, it replaces dense MLA with LongCat Sparse Attention (LSA) and natively supports context lengths of up to 1M tokens.
NV
NemotronLabs-VoiceChat
NVIDIA·🇺🇸·Major lab
Otheropen
▶ Hear it first. Natural turn-taking, barge-in and live tool calling.
1 variant11B
LA
K-EXAONE-2.0
LG AI Research·🇰🇷·Major lab
Textopen
We introduce K-EXAONE 2.0, a frontier-scale multilingual language model developed by LG AI Research. K-EXAONE 2.0 was scaled to more than three times the size of its predecessor through upcycling, followed by continual pretraining, difficulty-focused mid-training, and post-training.
We introduce Intern-S2-Mobius, a 35B foundation model built on the Mobius-v0 architecture realized by Xtuner and LMDeploy.
1 variantFP8
AT
UEmbed
Alibaba Tongyi Lab·🇨🇳·Research
Embeddingopen
UEmbed is a decoder-only multimodal embedding model that produces both dense embeddings and SPLADE-style sparse lexical embeddings from a single causal forward pass. It supports text, image, video, and mixed-modal inputs for retrieval, multimodal search, and visual-document retrieval.
3 variants9B4B2B
MI
MiniMax-H3
MiniMax·🇨🇳·Major lab
Videoopen
MiniMax H3 is a general-purpose video generation model with native stereo audio support. It takes text, images, videos, and audio as inputs and outputs synchronized audio-video clips between 4 and 15 seconds long at up to 2K resolution and 24 FPS. The input side is highly flexible, supporting first-and-last-frame mode, multi-reference inputs with up to nine images, three video clips, and three audio clips combined. A dedicated preprocessing component called H3-Context-IR parses and relates the multimodal context before generation. The model communicates fluently in eleven languages. Performance details beyond the specification sheet are not covered in the provided sources.
AQ
Qwen3.7 Flash
Alibaba Qwen·🇨🇳·Frontier
TextReasoningVisionAPI
Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...
TM
Inkling Small
Thinking Machines·🇺🇸·Major lab
TextReasoningVisionopenAPI
Inkling Small is a general-purpose multimodal model accepting text, image, and audio inputs to generate text. Its 276B-parameter sparse MoE architecture uses 12B active parameters per token, making it suited for agentic systems, coding assistance, and conversation.
2 variantsFreeNVFP4
Week of July 20, 202615 models
MI
Mage-ViT
Microsoft·🇺🇸·Major lab
Embeddingopen
Mage-ViT is the visual encoder at the core of Mage-VL. It is a Codec-ViT built primarily for video, where a single image is simply the degenerate one-frame case.
MI
Mage-VL
Microsoft·🇺🇸·Major lab
TextVisionopen
Mage-VL combines a from-scratch 4B visual encoder with a Qwen3-4B decoder to process images and video using codec-derived frame sparsity, cutting visual tokens by over 75%. Its dual-process design routes routine content through a lightweight gating mechanism while invoking the full model for event-worthy moments, enabling proactive streaming.
AN
Claude Opus 5
Anthropic·🇺🇸·Frontier
TextReasoningVisionAPI
Claude Opus 5.5 improves communication clarity over previous Opus models and achieves token efficiency approaching Fable 5.1 at lower cost. Output is capped at 128K tokens, making extended chain-of-thought unreliable.
1 variantBatch
MI
VibeVoice-ASR-BitNet
Microsoft·🇺🇸·Major lab
Audioopen
VibeVoice-ASR-BitNet is a compressed variant of VibeVoice-ASR optimized for real-time inference on edge CPUs — no GPU required.
SA
Apertus-v1.5
Swiss AI Initiative·🇨🇭·Research
TextVisionopen
No description yet — the model card or announcement hasn't been summarised.
2 variants70B8B
TE
Hy-MT2
Tencent·🇨🇳·Major lab
TextopenAPI
Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs, with workflows for structured, delimiter-based, contextual, glossary-based, and style-guided...
4 variants1.8B30B A3B7B30B A3B GGUF
AG
Ling 3.0 Flash
Ant Group (inclusionAI)·🇨🇳·Major lab
TextReasoningopenAPI
Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed with token efficiency and production-scale agentic inference as key priorities, enabling developers...
8 variantsSingprobeGGUFBaseBase 30TBase Midtrain
LI
LTX-2.5
Lightricks·🇮🇱·Major lab
Videoopen
No description yet — the model card or announcement hasn't been summarised.
1 variantDiffusers
KW
KAT-Coder-V2.5-Dev
Kwaipilot (Kuaishou)·🇨🇳·Major lab
TextCodeopen
This repository contains the model weights and configuration files for the post-trained KAT-Coder-V2.5-Dev in the Hugging Face Transformers format. The artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
AM
Instella-MoE
AMD·🇺🇸·Major lab
Textopen
Instella-MoE✨: Fully Open State-of-the-Art Mixture-of-Experts Language Model
Solar-Open2 is a 250 billion parameter large language model designed for agentic workflows, such as office productivity and coding. It utilizes a Hybrid-Attention Mixture-of-Experts architecture, allowing efficient inference even with long context lengths of up to 1 million tokens.
1 variant250B
GD
Gemini 3.6 Flash
Google DeepMind·🇺🇸·Frontier
TextReasoningVisionAPI
Gemini 3.6 Flash consumes 17 % fewer output tokens than its predecessor and completes multi-step workflows with fewer reasoning steps and tool calls. It targets coding, knowledge work, and agentic automation at lower cost.
1 variantBatch
GD
Gemini 3.5 Flash Lite
Google DeepMind·🇺🇸·Frontier
TextReasoningVisionAPI
Fastest and most cost-effective 3.5-class model delivering 350 output tokens per second according to the Artificial Analysis Index, significantly outperforming prior Flash-Lite generations in agentic workflows.
1 variantBatch
NV
Cosmos-H-Dreams
NVIDIA·🇺🇸·Major lab
Videoopen
Cosmos-H-Dreams is a real-time, action-conditioned generative surgical world model that lets a human operator or a learned surgical-robotics policy act inside a synthesized surgical scene and observe the interactions live.
MT
Motif-3
Motif Technologies·🇰🇷·Major lab
Textopen
Motif 3 Base is the base pretrained checkpoint of Motif 3 — a large-scale, decoder-only Mixture-of-Experts (MoE) language model with 314 billion total parameters and 13.2 billion parameters activated per token. It is built from the ground up by Motif Technologies following a fully in-house, proprietary design.
3 variantsBaseNVFP4Beta
Week of July 13, 202611 models
GD
Gemini 3.5 Flash Cyber
Google DeepMind·🇺🇸·Frontier
Text
A lightweight cybersecurity model built on Gemini 3.5 Flash and fine-tuned to find, validate, and patch software vulnerabilities in code more efficiently than the mainline Flash models. It is designed for use by security agents scanning large codebases at scale.
ME
Muse Spark 1.1
Meta·🇺🇸·Frontier
TextReasoningVisionAPI
Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...
MA
Shieldstral-1.0
Mistral AI·🇫🇷·Frontier
Otheropen
Shieldstral is a compact 3B-parameter, policy-adaptive multimodal safety classifier. Instead of predicting a fixed set of moderation categories, Shieldstral evaluates content against a safety policy expressed in natural language and returns a single continuous safety score.
1 variant3B
AG
LLaDA2.2-flash
Ant Group (inclusionAI)·🇨🇳·Major lab
Textopen
LLaDA2.2-flash is an agent-oriented diffusion language model in the LLaDA2 series.
PE
PII-Tracer
Perplexity·🇺🇸·Major lab
Embeddingopen
This repository provides PII-Tracer, the detector introduced in PII-TRACE.
2 variantsMLXGGUF
SA
Intern-S2
Shanghai AI Lab (InternLM)·🇨🇳·Research
TextVisionopen
Intern-S2 is a multimodal foundation model for scientific intelligence and long-horizon agentic tasks. It unifies vision-language pre-training from raw scientific literature pages with large-scale multi-task reinforcement learning across 20+ domains, plus black-box agentic RL in sandboxed environments. Tool calling, images, video, and time-series inputs are supported.
No description yet — the model card or announcement hasn't been summarised.
3 variants14B Distilled Diffusers14B Diffusers14B
TM
Inkling
Thinking Machines·🇺🇸·Major lab
TextReasoningVisionopenAPI
Inkling ist ein allgemeines Multimodell von Thinking Machines mit 975 Milliarden Parametern (41 Milliarden aktiv), das Text, Bilder und Audio versteht und Text generiert. Es nutzt eine MoE-Architektur mit kontrollierbarer Reasoning-Effort und ist auf agentic Workflows, Coding und wissenschaftliches Reasoning ausgelegt.
2 variantsFreeNVFP4
TE
Hy-Embodied-VLM-1.0
Tencent·🇨🇳·Major lab
TextVisionopen
Hy-Embodied-VLM-1.0 Efficient Physical-World Agents Tencent Robotics X × Hy Vision Team × Futian Laboratory
TE
Hy-Embodied-RxBrain-1.0
Tencent·🇨🇳·Major lab
TextVisionopen
RxBrain Embodied Cognition Foundation Model with Joint Language–Visual Reasoning and Imagination Tencent Robotics X × Futian Laboratory × Tencent Hy Team
NV
Ising-Calibration-1.5
NVIDIA·🇺🇸·Major lab
TextVisionopen
NVIDIA-Ising-Calibration-1.5-31B-BF16 is a dense multimodal vision-language model built on Gemma 4 31B.
2 variants31B NVFP431B BF16
Week of July 6, 202614 models
AQ
Wan-Dancer
Alibaba Qwen·🇨🇳·Frontier
Videoopen
Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation
1 variant14B
KW
KAT-Coder-Pro V2.5
Kwaipilot (Kuaishou)·🇨🇳·Major lab
TextCodeAPI
KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...
OP
GPT-5.6 Luna Pro
OpenAI·🇺🇸·Frontier
TextReasoningVisionAPI
GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoningreasoning-mode
1 variantBatch
Same release
XA
Grok 4.5
xAI·🇺🇸·Frontier
TextReasoningVisionAPI
Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.
AI
Aion-3.0-Mini
AionLabs·🇮🇱·Major lab
TextReasoningAPI
Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each...
Same release
NV
Cosmos3-Super-Image2Video-4Step
NVIDIA·🇺🇸·Major lab
Videoopen
NVIDIA Cosmos™ is a world foundation model platform designed to accelerate the development of Physical AI by enabling machines to understand, simulate, and interact with the physical world across robotics, autonomous driving, and smart space environments, including industrial and factory-scale applications.
NV
Cosmos3-Super-Text2Image-4Step
NVIDIA·🇺🇸·Major lab
Imageopen
NVIDIA Cosmos™ is a world foundation model platform designed to accelerate the development of Physical AI by enabling machines to understand, simulate, and interact with the physical world across robotics, autonomous driving, and smart space environments, including industrial and factory-scale applications.
NV
Nemotron-Labs-Audex
NVIDIA·🇺🇸·Major lab
Textopen
We're excited to introduce Nemotron-Labs-Audex-2B, a unified audio-text LLM with a similar recipe as Nemotron-Labs-Audex-30B-A3B. Audex-2B extends the vocabulary for discrete audio tokens used for speech and general audio outputs, as well as an audio encoder for speech and general audio inputs.
2 variants2B30B A3B
Week of June 29, 202610 models
ML
LongCat 2.0
Meituan LongCat·🇨🇳·Major lab
TextReasoningopenAPI
LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding, repository-level changes, long-horizon problem solving, and agentic...
2 variantsINT8FP8
BS
PAR
ByteDance Seed·🇨🇳·Major lab
Otheropen
Three 3-scale PAR model variants are provided.
PO
Laguna S 2.1
Poolside·🇺🇸·Major lab
TextReasoningopenAPI
Laguna S 2.1 is the latest coding agent model from Poolside. Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...
6 variantsFreeNVFP4 MLXGGUFINT4NVFP4
MA
Leanstral-1.5
Mistral AI·🇫🇷·Frontier
Otheropen
Leanstral 1.5 is an open-source code agent model designed for Lean 4, a proof assistant capable of expressing complex mathematical objects such as perfectoid spaces and software specifications like properties of Rust fragments.
1 variant119B A6B
NV
Cosmos3-Edge
NVIDIA·🇺🇸·Major lab
Otheropen
NVIDIA Cosmos3-Edge is a 4B-parameter omni-modal world model that generates text, images, video, and action commands from multimodal inputs. Its Mixture-of-Transformers architecture combines autoregressive decoding for text with diffusion-based denoising for other modalities. Optimized for physical AI, it serves robotics, autonomous vehicles, and smart space simulations.
AN
Claude Sonnet 5
Anthropic·🇺🇸·Frontier
TextReasoningVisionAPI
Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...
1 variantBatch
GD
Nano Banana 2 Lite
Google DeepMind·🇺🇸·Frontier
TextReasoningVisionImageAPI
A text-to-image generation and editing model used for creating graphics, posters, and social-media content. It supports object segmentation, direct in-image text editing, and is integrated into Google Workspace apps.
NV
Nemotron-Parse-2.0
NVIDIA·🇺🇸·Major lab
TextVisionopen
NVIDIA Nemotron Parse 2.0 transforms document images into structured, machine-readable representations with text, layout classes, bounding boxes, and reading-order information.
MO
Moondream3.1
Moondream·🇺🇸·Major lab
TextVisionopen
Moondream 3.1 is a vision language model with a mixture-of-experts architecture (9B total parameters, 2B active). It delivers state-of-the-art visual reasoning and detection while staying fast and cheap to deploy.
1 variant9B A2B
GD
Tabfm 1.0.0
Google DeepMind·🇺🇸·Frontier
Otheropen
TabFM is a zero-shot tabular foundation model from Google Research.
2 variantsJaxPytorch
Week of June 22, 20264 models
AI
OlmoEarth-v1_2
Allen Institute for AI (Ai2)·🇺🇸·Research
Embeddingopen
OlmoEarth-v11-Base is a ViT-Base (114M parameters) foundation model for remote sensing tasks over Sentinel-1, Sentinel-2, and Landsat images or image time series.
1 variantBase
SA
Fugu Ultra
◆ Underrated
Sakana AI·🇯🇵·Major lab
TextVision
Sakana AI's higher-performance Fugu model — not a monolithic LLM but a learned multi-agent orchestration system that routes requests across multiple models and tools. Multimodal (text+image input), 1M context, $5/$30 per MTok on OpenRouter.
NV
Nemotron-Labs-3-Puzzle
NVIDIA·🇺🇸·Major lab
Textopen
Nemotron-Labs-3-Puzzle-75B-A9B is a deployment-optimized large language model developed by NVIDIA, derived from Nemotron-3-Super-120B-A12B.
3 variants75B A9B FP875B A9B BF1675B A9B NVFP4
AQ
AgentWorld
Alibaba Qwen·🇨🇳·Frontier
Textopen
Qwen-AgentWorld is the first language world model to cover seven agent interaction domains within a single model. It simulates agentic environments via long chain-of-thought reasoning, predicting the next environment state given an agent's action and interaction history.
1 variant35B A3B
Week of June 15, 20269 models
ZA
GLM-5.2
◆ Underrated
Z.ai (Zhipu)·🇨🇳·Frontier
TextReasoningopenAPI
GLM-5.2 is an open-source model from Z.AI for long-horizon tasks, with a stable 1-million-token context and strong agentic coding. It combines code generation with tool use and supports adjustable thinking-effort levels. It scores 62.1% on SWE-bench Pro, sits close to the frontier on AIME and GPQA-Diamond, and improves markedly over GLM-5.1. The MoE architecture with IndexShare cuts compute per token by a factor of 2.9.
1 variantFP8
PO
Laguna 2.1
Poolside·🇺🇸·Major lab
TextReasoningopenAPI
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in April 2026). It combines...
Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.3 + CUDA12.9:
GD
Nano Banana Pro
Google DeepMind·🇺🇸·Frontier
TextReasoningVisionImageAPI
Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...
CO
Transcribe Arabic
Cohere·🇨🇦·Major lab
Audioopen
No description yet — the model card or announcement hasn't been summarised.
1 variant202607
GD
Gemini 3.1 Flash Image
Google DeepMind·🇺🇸·Frontier
VisionImageAPI
Google's latest Gemini Flash generation with native image input/output. Available on OpenRouter ($0.50/$3 per MTok). Gemini 3.x is a standalone series alongside Gemini 2.5.
MI
FastContext 1.0
◆ Underrated
Microsoft·🇺🇸·Major lab
Codeopen
Microsoft's specialized code-search model for coding agents — powers the 'Explore' subagent in SWE-FastContext. Not a general-purpose LLM, but relevant for agentic setups.
1 variant4B
CO
North Mini Code 1.0
◆ Underrated
Cohere·🇨🇦·Major lab
CodeopenAPI
North Mini Code is an open-weight sparse MoE model (30B total, 3B active) specialized for code generation, agentic software engineering, and terminal tasks. It supports tool-use through chat templates and was post-trained with supervised fine-tuning followed by reinforcement learning with verifiable rewards focused on agentic coding.
2 variantsW4A16FP8
AI
MolmoMotion-H3-F30
Allen Institute for AI (Ai2)·🇺🇸·Research
TextVisionopen
MolmoMotion is a 4B vision-language model that forecasts 3D point trajectories under natural-language action instructions.
1 variant4B
Week of June 8, 202614 models
MA
Kimi K3
Moonshot AI·🇨🇳·Frontier
TextReasoningVisionopenAPI
Kimi K3 ist ein Open-Weight-Multimodalmodell von Moonshot AI mit 2,8 Billionen Parametern (davon 104 Mrd. aktiv) im MoE-Design und nativer Bild- sowie Videoerkennung. Die auf Kimi Delta Attention und Attention Residuals aufbauende Architektur zielt auf langfristiges Coding, Reasoning und agentic Knowledge Work ab; das Modell unterstuetzt einen Kontextfenster von 1 Million Tokens und versteht Text, Bilder und Videos im selben Modell.
1 variantBatch
MA
Kimi K2.7 Code
Moonshot AI·🇨🇳·Frontier
ReasoningCodeopenAPI
Kimi K2.7 Code is Moonshot AI's coding-specialised agentic model. It accepts image input and tool calls and can handle complex software-engineering workflows end to end. It scores 62.0 on Kimi Code Bench v2 and uses about 30% fewer thinking tokens than its predecessor K2.6. It trails GPT-5.5 slightly on coding benchmarks but beats Claude Opus 4.8 on MCP Mark Verified with 81.1.
AG
VISTA
Ant Group (inclusionAI)·🇨🇳·Major lab
TextVisionopen
VISTA-9B are GUI-grounding vision-language models trained from Qwen3.5 9B backbones with VISTA: View-Consistent Self-Verified Training for GUI Grounding.
2 variants9B4B
LA
LFM2.5-Encoder
Liquid AI·🇺🇸·Major lab
TextEmbeddingopen
LFM2.5-Encoder is a family of multilingual bidirectional encoders built on the LFM2 architecture, available in two sizes:
No description yet — the model card or announcement hasn't been summarised.
2 variantsBaseSFT
ZY
ZONOS2
Zyphra·🇺🇸·Major lab
Audioopen
ZONOS2 is our latest text-to-speech model trained on more than 6 million hours of varied multilingual speech, delivering expressiveness and quality on par with—or even surpassing—top TTS providers at low latency with MoE. ZONOS2 excels at high-fidelity and naturalistic voice cloning.
1 variantGGUF
BS
EvoQuality
ByteDance Seed·🇨🇳·Major lab
TextVisionopen
- Model Name: EvoQuality (Self-Evolving VLM for Image Quality Assessment) - Task: No-Reference Image Quality Assessment (NR-IQA), supporting both single-image quality scoring and pairwise quality comparison (ranking) - Core Idea: Without relying on any human-annotated quality scores or distortion-type labels,…
OP
UltraX
OpenBMB·🇨🇳·Research
Textopen
UltraX is a function-calling refinement framework for large-scale pre-training data.
1 variant0.6B Preview
AN
Claude Fable 5
Anthropic·🇺🇸·Frontier
ReasoningVisionAPI
Anthropic's new top flagship — replaces the Opus tier as the strongest generally available model. Arena Elo 1508 (rank 1). 1M context, Adaptive Thinking always on, $10/$50 per MTok. Fable 5 + Mythos 5 (invite-only) launched together on June 9, 2026.
GD
Diffusiongemma
Google DeepMind·🇺🇸·Frontier
TextVisionopen
DiffusionGemma is a generative model built by Google DeepMind. Based on the 26B A4B Mixture-of-Experts (MoE) Gemma 4 architecture, DiffusionGemma generates tokens using discrete diffusion.
1 variant26B A4B IT
ZA
SCAIL-2
Z.ai (Zhipu)·🇨🇳·Frontier
Videoopen
SCAIL-2 is an open-source model for end-to-end controlled character animation. It animates a reference character with a driving video, and also supports character replacement and multi-character scenarios without relying on intermediate pose representations.
BA
PP-OCRv6_tiny
Baidu·🇨🇳·Major lab
TextVisionopen
PP-OCRv6: From 1.5M to 34.5M Parameters, Surpassing Billion-Scale VLMs on OCR Tasks
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...
2 variantsSafetySafety Free
NV
ArtiFixer
NVIDIA·🇺🇸·Major lab
Otheropen
ArtiFixer is a few-step causal auto-regressive model that enhances and extends 3D reconstruction. The related source code provides implementations for training, evaluation, and inference, supporting various stages including bidirectional training, diffusion forcing, and Self-Forcing-style DMD distillation.
NA
Nex-N2-Pro
◆ Underrated
Nex AGI (Shanghai Innovation Inst.)·🇨🇳
ReasoningTextopenAPI
397B model from the Shanghai Innovation Institute — open source, Apache 2.0, free on OpenRouter. Agentic-focused, with papers on a 'Unified Ecosystem for Large-Scale Environment Construction'. 7.87K HF likes despite barely any Western coverage.
AQ
Qwen3.7 Plus
Alibaba Qwen·🇨🇳·Frontier
ReasoningTextAPI
Qwen's latest Plus tier — 1M context, $0.32/$1.28 per MTok. Qwen3.7 Max (stronger) and Plus (more efficient) form the top of the Qwen3.7 line.
LA
LFM2.5
Liquid AI·🇺🇸·Major lab
TextReasoningopenAPI
LFM2.5 is a hybrid model designed for on-device deployment, featuring 230 million parameters and optimized for fast edge inference. It is suitable for general-purpose text generation and agentic tasks, but not recommended for reasoning-heavy workloads.
MiniMax is barely known in the West — but its M3 model on OpenRouter ($0.30/$1.20 per MTok) is one of the cheapest 1M-context multimodal providers. Successor to Text-01 (4M context).
1 variantMXFP8
KY
Moshika
Kyutai·🇫🇷·Research
Audioopen
No description yet — the model card or announcement hasn't been summarised.
1 variantRL Seamless
DE
DeepSeek V4 Pro
DeepSeek·🇨🇳·Frontier
ReasoningTextopenAPI
DeepSeek's heavyweight: 1.6T total, 49B active, 1M context, FP4/FP8 mixed precision. Codeforces 3206 Elo beats GPT-5.4. MIT license, fully open weights. $0.435/$0.87 per MTok on OpenRouter.
3 variants081308130813 Batch
Same release
BS
Bernini-R
ByteDance Seed·🇨🇳·Major lab
Otheropen
- [2026-06-09] We open-sourced the 1.3B weights of the Bernini Renderer (Bernini-R) on ByteDance/Bernini-R-1.3B-Diffusers.
2 variants1.3B DiffusersDiffusers
Week of May 25, 20262 models
AN
Claude Opus 4.8
Anthropic·🇺🇸·Frontier
ReasoningVisionAPI
Current Opus-tier model (until Claude Fable 5). Arena Elo ~1490. 1M context, Adaptive Thinking, knowledge through Jan 2026. $5/$25 per MTok direct, similar on OpenRouter.
ST
Step 3.7 Flash
◆ Underrated
StepFun·🇨🇳·Major lab
VisionReasoningopenAPI
201B multimodal open-weights model from StepFun — underrated in the West. Image+text input, strong reasoning. Successor to Step 3.5 Flash (199B). $0.20/$1.15 per MTok on OpenRouter.
Week of May 18, 20264 models
AQ
Qwen3.7 Max
Alibaba Qwen·🇨🇳·Frontier
ReasoningTextAPI
Strongest Qwen3.7 variant: 1M context, $1.25/$3.75 per MTok. More capable than Plus, but pricier. The Qwen3.7 line shipped late May 2026 as the successor to Qwen3.6.
XA
Grok Build 0.1
◆ Underrated
xAI·🇺🇸·Frontier
CodeReasoningAPI
xAI's coding-focused model — 256K context, $1/$2 per MTok. 'Build' implies software development as the main use case. A separate coding line alongside the general-purpose Grok 4.3.
GD
Gemini 3.5 Flash
Google DeepMind·🇺🇸·Frontier
ReasoningVisionAPI
Google's current Flash generation: 1M context, all modalities, $1.50/$9 per MTok on OpenRouter. Gemini 3.5 Flash is the speed tier of the Gemini 3.x series that replaces Gemini 2.5.
MA
Kimi K2.6
Moonshot AI·🇨🇳·Frontier
ReasoningVisionopenAPI
Multimodal predecessor of K2.7-Code — Visual Agentic Intelligence with 1.1T parameters. 2.66M HF downloads. Still relevant for multimodal tasks (images, video) since K2.7 is code-only.
Week of May 11, 20261 model
ZA
GLM-5.1
◆ Underrated
Z.ai (Zhipu)·🇨🇳·Frontier
TextReasoningopen
Direct predecessor of GLM-5.2 (June 20). Still interesting for local deployment. The GLM-5.1 FP8 variant has 1.33M downloads — massive interest from the CN community.
Week of May 4, 20261 model
GD
Gemini 3.1 Flash Lite
Google DeepMind·🇺🇸·Frontier
VisionTextAPI
Google's cheapest Gemini 3.x option: 1M context, $0.25/$1.50 per MTok. The 'Lite' variant is the first choice when token costs are critical and multimodality is needed.
Week of April 27, 202611 models
PO
Laguna M.1
◆ Underrated
Poolside·🇺🇸·Major lab
CodeReasoningopenAPI
Poolside's strong coding-agent model: 226B MoE, Apache 2.0, SWE-bench Pro 49.2%. Free on OpenRouter! Specialized for agentic long-horizon coding tasks. Poolside is a little-known US AI company.
1 variantBase
PO
Laguna XS.2
◆ Underrated
Poolside·🇺🇸·Major lab
CodeopenAPI
Poolside's lean coding model: 33B MoE, 3B active, for local deployments. Free on OpenRouter. Sliding-window attention for very fast inference. 231K downloads on HF.
1 variantXS
XA
Grok 4.3
xAI·🇺🇸·Frontier
ReasoningVisionAPI
xAI's current frontier model on OpenRouter ($1.25/$2.50 per MTok). Grok 4.3 positions itself as a strong all-rounder. Also accessible via Grok.com / X Premium.
MA
Mistral Small 4
Mistral AI·🇫🇷·Frontier
TextReasoningopenAPI
Mistral ironically calls 119B 'Small' — Apache 2.0, with instruction-following, reasoning and coding in one. Mistral positions itself as the strongest European open-weights challenger.
1 variant119B
MA
Mistral Medium 3.5
Mistral AI·🇫🇷·Frontier
TextVisionAPI
Mistral Medium 3.5 is Mistral's first flagship to handle instruction-following, reasoning and coding in a unified way. 418K HF downloads, EAGLE-acceleration variant available. $1.50/$7.50 per MTok on OpenRouter.
IB
Granite 4.1
◆ Underrated
IBM·🇺🇸·Major lab
TextCodeopenAPI
IBM's Granite 4.x generation: 8B, Apache 2.0, $0.05/$0.10 per MTok (one of the cheapest). Enterprise-focused, strong on structured tasks. IBM ships Granite models consistently with open weights.
1 variant8B
MA
Kimi K2.5
Moonshot AI·🇨🇳·Frontier
ReasoningVisionopen
Moonshot's 'Visual Agentic Intelligence' — 1.81M HF downloads. Basis for K2.6 and K2.7. One of the first large models with real multimodal agentics at 1T parameters.
CO
Command A+
◆ Underrated
Cohere·🇨🇦·Major lab
ReasoningVisionopenAPI
Cohere's first multimodal open-weights flagship: 218B MoE, 25B active, 128K context, 48 languages, Apache 2.0. Reasoning tokens, tool use with JSON schema. Not on OpenRouter yet.
NV
Nemotron 3 Nano Omni
◆ Underrated
NVIDIA·🇺🇸·Major lab
ReasoningVisionopenAPI
NVIDIA's small multimodal reasoning model: 30B MoE, only 3B active, text+image+audio input. Free on OpenRouter. Designed for on-device, 4x faster than its predecessor, reasoning ON/OFF mode.
1 variant30B
AQ
Qwen3.6 Flash
Alibaba Qwen·🇨🇳·Frontier
ReasoningTextAPI
Fast variant of the Qwen3.6 family: 1M context, $0.19/$1.13 per MTok. Qwen3.6 shipped as Flash, 27B, 35B and Max-preview variants — all on April 28, 2026.
Same release
Week of April 20, 20268 models
OP
GPT-5.5
OpenAI·🇺🇸·Frontier
ReasoningVisionAPI
OpenAI's current frontier model: 'A new class of intelligence for coding and professional work.' 1M context, knowledge cutoff Dec 2025, $5/$30 per MTok. GPT-5.5 Pro as the stronger variant ($30/$180 per MTok).
Same release
XM
MiMo V2.5 Pro
◆ Underrated
Xiaomi MiMo·🇨🇳·Major lab
ReasoningTextopenAPI
Xiaomi's reasoning model on OpenRouter ($0.435/$0.87 per MTok). MiMo is Xiaomi's first serious LLM push — barely noticed in the West. 1M context. V2.5 (lighter) and V2.5-Pro available.
1 variantFP4 Dflash
Same release
TE
Hy3
◆ Underrated
Tencent·🇨🇳·Major lab
TextReasoningopenAPI
Tencent's Hunyuan 3 in preview on OpenRouter ($0.063/$0.21 per MTok). One of the cheapest models available. Tencent is a heavyweight that gets little Western attention in AI.
2 variantsPreviewFP8
AG
Ling-2.6
◆ Underrated
Ant Group (inclusionAI)·🇨🇳·Major lab
ReasoningTextopenAPI
Ant Group's (Alipay's parent) 1-trillion-parameter MoE model — Apache 2.0. Extremely cheap on OpenRouter ($0.075/$0.625 per MTok). 472 likes on HF despite barely any Western coverage.
2 variants1T1T Base
Same release
OP
GPT-5.4 Image 2
OpenAI·🇺🇸·Frontier
ImageVisionAPI
OpenAI's second image-generation iteration on the GPT-5.4 base: $8/$15 per MTok (input/output). The current standard route for professional AI image generation via OpenRouter API.
Week of March 30, 20261 model
ZA
GLM-5
◆ Underrated
Z.ai (Zhipu)·🇨🇳·Frontier
TextReasoningopen
First release of the GLM-5 generation (April 2026). 2.1K likes on HuggingFace. Zhipu's development rhythm is remarkable: one major release per month.
Week of February 23, 20261 model
GD
Gemma 3n
◆ Underrated
Google DeepMind·🇺🇸·Frontier
Visionopen
Google's on-device multimodal model: handles text, image, video AND audio at just 4B effective parameters. MatFormer architecture enables sub-models within the same checkpoint. Designed for edge deployments.
1 variantE4B
Week of January 26, 20261 model
AN
Claude Opus 4.7
Anthropic·🇺🇸·Frontier
ReasoningVisionAPI
A few generations before Fable 5 — but in thinking mode Arena Elo 1502, rank 3 worldwide. Still relevant for setups where Extended Thinking is desired and Fable 5's Adaptive Thinking isn't enough.
Week of October 27, 20251 model
OP
GPT-5.4
OpenAI·🇺🇸·Frontier
ReasoningVisionAPI
Predecessor of GPT-5.5, but still very widely used ($2.50/$15 per MTok vs. $5/$30 for 5.5). Mini variant (400K context, $0.75/$4.50) for coding agents and subagents. IMOAnswerBench 91.4%.
Week of April 28, 20251 model
AQ
Qwen3
◆ Underrated
Alibaba Qwen·🇨🇳·Frontier
ReasoningTextopen
Alibaba's strong open-weights model from April 2025 — basis for many downstream variants. Qwen3.5 and 3.6 followed in 2026. Often overlooked in the West despite AIME 2025 85.7%.
1 variant235B A22B
Week of March 31, 20251 model
ME
Llama 4 Maverick
Meta·🇺🇸·Frontier
Visionopen
Meta's multimodal MoE flagship: 17B active, 128 experts, 402B total, 1M-token context. Scout variant (16E, 109B) for lean deployments. Llama 5 is expected in H2 2026.
Week of March 10, 20251 model
GD
Gemma 3
Google DeepMind·🇺🇸·Frontier
Visionopen
Google's open-weights flagship: 27B (also 1B, 4B, 12B), 128K context, text+image, 140+ languages. Gemma 3 is the basis for many community fine-tunes. Superseded by Gemma 4.