Ling-3.0-Flash-VL
Ling-3.0-Flash-VL extends the Ling3.0-Flash language backbone with native image and video understanding for visual reasoning and agent execution. A ViT encoder, MLP projector, VideoRoPE, and hybrid KDA/gated MLA backbone combine visual evidence with planning, action, and verification. The 124B model activates 5.5B parameters and supports up to 256K context through the documented deployment extension. The September 4 official changelog establishes public chat/API access before the subsequent open-weight promotion. MIT weights enable local visual assistants and document or video analysis.
2026-09-04
124B total, 5.5B active
Vision encoder plus hybrid KDA/Gated MLA Mixture of Experts
MIT
Specifications
- Parameters
- 124B total, 5.5B active
- Architecture
- Vision encoder plus hybrid KDA/Gated MLA Mixture of Experts
- License
- MIT
- Context Window
- 262,144 tokens
- Type
- multimodal
- Modalities
- textimagevideo
Benchmark Scores
Advanced Specifications
- Model Family
- Ling
- Finetuned From
- Ling-3.0-flash
- API Access
- Available
- Chat Interface
- Available
Capabilities & Limitations
- Capabilities
- reasoningcodingtool useinstruction followingvisual reasoningvideo understandingGUI interaction
- Known Limitations
- Inactive experts still require substantial weight storage.Tool use requires an external execution environment and appropriate chat template.256K operation uses the documented YaRN deployment recipe.
- Notable Use Cases
- coding agentsenterprise assistantslong-document analysisvisual agents
- Function Calling Support
- Yes
- Tool Use Support
- Yes