Ling-2.6-Flash
Ling-2.6-Flash is Ant Group’s efficient instruction-following tier, designed for faster responses and lower token overhead in everyday agents. The 104B MoE network activates 7.4B parameters, combining MLA and Lightning Linear Attention in a 1:7 ratio. Post-training focuses on concise execution, coding, planning, and tool use rather than long visible reasoning. The first-party April 22 changelog documents public API and chat availability, preceding the later weight upload. MIT weights enable local serving, with a native 128K context and demonstrated YaRN extension to 256K.
2026-04-22
104B total, 7.4B active
Hybrid MLA/Lightning Linear Attention Mixture of Experts
MIT
Specifications
- Parameters
- 104B total, 7.4B active
- Architecture
- Hybrid MLA/Lightning Linear Attention Mixture of Experts
- License
- MIT
- Context Window
- 131,072 tokens
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- Ling
- API Access
- Available
- Chat Interface
- Available
Capabilities & Limitations
- Capabilities
- reasoningcodingtool useinstruction following
- Known Limitations
- Inactive experts still require substantial weight storage.Tool use requires an external execution environment and appropriate chat template.
- Notable Use Cases
- coding agentsenterprise assistantslong-document analysis
- Function Calling Support
- Yes
- Tool Use Support
- Yes