Ling-3.0-Flash
Ling-3.0-Flash is a native hybrid reasoning model using Kimi Delta Attention and gated Multi-head Latent Attention in a 5:1 layer pattern. Sparse experts support coding, knowledge work, mathematical reasoning, and tool-based agents, with thinking configurable per request. The Flash tier targets production agent efficiency with 124B total parameters. Public API availability precedes the later August weight publication, as established by the official changelog. MIT weights, routine base checkpoints, and quantized distributions are grouped under this tier.
2026-07-23
124B total, 5.1B active
Hybrid KDA/Gated MLA Mixture of Experts
MIT
Specifications
- Parameters
- 124B total, 5.1B active
- Architecture
- Hybrid KDA/Gated MLA Mixture of Experts
- License
- MIT
- Context Window
- 131,072 tokens
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- Ling
- API Access
- Available
- Chat Interface
- Available
- Variants
- Ling-3.0-flash-baseBF16/FP8/INT4 weight distributions
Capabilities & Limitations
- Capabilities
- reasoningcodingtool useinstruction following
- Known Limitations
- Inactive experts still require substantial weight storage.Tool use requires an external execution environment and appropriate chat template.
- Notable Use Cases
- coding agentsenterprise assistantslong-document analysis
- Function Calling Support
- Yes
- Tool Use Support
- Yes