Gander
Gander is a 9B full-duplex audio-visual interaction model developed by Tencent’s Hunyuan Speech team with academic collaborators. Its MiniCPM-o 4.5-derived Thinker and streaming Talker handle speech, images, video frames, text, turn taking, interruption, and structured delegation while tasks continue asynchronously. Apache-licensed model components and a serving runtime are publicly available. The larger system pairs this conversational model with a separately configured reasoning agent for long-horizon execution; those back-end model capabilities are not assumed to reside inside the 9B checkpoint. The release date follows the September 8 technical report that publishes the named model; the project dates its checkpoint release to September 9.
Specifications
- Parameters
- 9B
- Architecture
- Multimodal Transformer with streaming Thinker–Talker architecture
- License
- Apache-2.0
- Type
- multimodal
- Modalities
- textimageaudiovideo
Benchmark Scores
Advanced Specifications
- Model Family
- Gander
- Finetuned From
- MiniCPM-o 4.5
- API Access
- Not Available
- Chat Interface
- Not Available
- Variants
- Gander-Unit8released Thinker and Talker pair
Capabilities & Limitations
- Capabilities
- full-duplex conversationstreaming speechaudio-visual understandinginterruption handlingtask delegation
- Known Limitations
- Long-horizon task execution depends on a configured external reasoning agentStreaming temporal units are not a standard token context limitPerception and factual errors remain possible
- Notable Use Cases
- realtime voice assistantsscreen-grounded conversational agentsmultimodal interaction research
- Tool Use Support
- Yes