Ling-3.0-Tiny
Ling-3.0-Tiny is a native hybrid reasoning model using Kimi Delta Attention and gated Multi-head Latent Attention in a 3:1 layer pattern. Sparse experts support coding, knowledge work, mathematical reasoning, and tool-based agents, with thinking configurable per request. The Tiny tier targets local and edge use with 7.9B total parameters; the card documents Apple Silicon and DGX Spark deployment. Public API availability precedes the later August weight publication, as established by the official changelog. MIT weights, routine base checkpoints, and quantized distributions are grouped under this tier.
2026-08-07
7.9B total, 1.3B active
Hybrid KDA/Gated MLA Mixture of Experts
MIT
Specifications
- Parameters
- 7.9B total, 1.3B active
- Architecture
- Hybrid KDA/Gated MLA Mixture of Experts
- License
- MIT
- Context Window
- 131,072 tokens
- Type
- text
- Modalities
- text
Benchmark Scores
Advanced Specifications
- Model Family
- Ling
- API Access
- Available
- Chat Interface
- Available
- Variants
- Ling-3.0-tiny-baseBF16/FP8/INT4 weight distributions
Capabilities & Limitations
- Capabilities
- reasoningcodingtool useinstruction following
- Known Limitations
- Inactive experts still require substantial weight storage.Tool use requires an external execution environment and appropriate chat template.
- Notable Use Cases
- coding agentsenterprise assistantslong-document analysis
- Function Calling Support
- Yes
- Tool Use Support
- Yes