InclusionAI (Ant Group) logo

Ling-2.6-Flash

InclusionAI (Ant Group)Open WeightsPending Human Review

Ling-2.6-Flash is Ant Group’s efficient instruction-following tier, designed for faster responses and lower token overhead in everyday agents. The 104B MoE network activates 7.4B parameters, combining MLA and Lightning Linear Attention in a 1:7 ratio. Post-training focuses on concise execution, coding, planning, and tool use rather than long visible reasoning. The first-party April 22 changelog documents public API and chat availability, preceding the later weight upload. MIT weights enable local serving, with a native 128K context and demonstrated YaRN extension to 256K.

2026-04-22
104B total, 7.4B active
Hybrid MLA/Lightning Linear Attention Mixture of Experts
MIT

Specifications

Parameters
104B total, 7.4B active
Architecture
Hybrid MLA/Lightning Linear Attention Mixture of Experts
License
MIT
Context Window
131,072 tokens
Type
text
Modalities
text

Benchmark Scores

Advanced Specifications

Model Family
Ling
API Access
Available
Chat Interface
Available

Capabilities & Limitations

Capabilities
reasoningcodingtool useinstruction following
Known Limitations
Inactive experts still require substantial weight storage.Tool use requires an external execution environment and appropriate chat template.
Notable Use Cases
coding agentsenterprise assistantslong-document analysis
Function Calling Support
Yes
Tool Use Support
Yes

Related Models