InclusionAI (Ant Group) logo

Ling-3.0-Flash-VL

InclusionAI (Ant Group)Open WeightsPending Human Review

Ling-3.0-Flash-VL extends the Ling3.0-Flash language backbone with native image and video understanding for visual reasoning and agent execution. A ViT encoder, MLP projector, VideoRoPE, and hybrid KDA/gated MLA backbone combine visual evidence with planning, action, and verification. The 124B model activates 5.5B parameters and supports up to 256K context through the documented deployment extension. The September 4 official changelog establishes public chat/API access before the subsequent open-weight promotion. MIT weights enable local visual assistants and document or video analysis.

2026-09-04
124B total, 5.5B active
Vision encoder plus hybrid KDA/Gated MLA Mixture of Experts
MIT

Specifications

Parameters
124B total, 5.5B active
Architecture
Vision encoder plus hybrid KDA/Gated MLA Mixture of Experts
License
MIT
Context Window
262,144 tokens
Type
multimodal
Modalities
textimagevideo

Benchmark Scores

Advanced Specifications

Model Family
Ling
Finetuned From
Ling-3.0-flash
API Access
Available
Chat Interface
Available

Capabilities & Limitations

Capabilities
reasoningcodingtool useinstruction followingvisual reasoningvideo understandingGUI interaction
Known Limitations
Inactive experts still require substantial weight storage.Tool use requires an external execution environment and appropriate chat template.256K operation uses the documented YaRN deployment recipe.
Notable Use Cases
coding agentsenterprise assistantslong-document analysisvisual agents
Function Calling Support
Yes
Tool Use Support
Yes

Related Models