InclusionAI (Ant Group) logo

Ling-3.0-Flash

InclusionAI (Ant Group)Open WeightsPending Human Review

Ling-3.0-Flash is a native hybrid reasoning model using Kimi Delta Attention and gated Multi-head Latent Attention in a 5:1 layer pattern. Sparse experts support coding, knowledge work, mathematical reasoning, and tool-based agents, with thinking configurable per request. The Flash tier targets production agent efficiency with 124B total parameters. Public API availability precedes the later August weight publication, as established by the official changelog. MIT weights, routine base checkpoints, and quantized distributions are grouped under this tier.

2026-07-23
124B total, 5.1B active
Hybrid KDA/Gated MLA Mixture of Experts
MIT

Specifications

Parameters
124B total, 5.1B active
Architecture
Hybrid KDA/Gated MLA Mixture of Experts
License
MIT
Context Window
131,072 tokens
Type
text
Modalities
text

Benchmark Scores

Advanced Specifications

Model Family
Ling
API Access
Available
Chat Interface
Available
Variants
Ling-3.0-flash-baseBF16/FP8/INT4 weight distributions

Capabilities & Limitations

Capabilities
reasoningcodingtool useinstruction following
Known Limitations
Inactive experts still require substantial weight storage.Tool use requires an external execution environment and appropriate chat template.
Notable Use Cases
coding agentsenterprise assistantslong-document analysis
Function Calling Support
Yes
Tool Use Support
Yes

Related Models