Tencent logo

Gander

TencentOpen WeightsPending Human Review

Gander is a 9B full-duplex audio-visual interaction model developed by Tencent’s Hunyuan Speech team with academic collaborators. Its MiniCPM-o 4.5-derived Thinker and streaming Talker handle speech, images, video frames, text, turn taking, interruption, and structured delegation while tasks continue asynchronously. Apache-licensed model components and a serving runtime are publicly available. The larger system pairs this conversational model with a separately configured reasoning agent for long-horizon execution; those back-end model capabilities are not assumed to reside inside the 9B checkpoint. The release date follows the September 8 technical report that publishes the named model; the project dates its checkpoint release to September 9.

2026-09-08
9B
Multimodal Transformer with streaming Thinker–Talker architecture
Apache-2.0

Specifications

Parameters
9B
Architecture
Multimodal Transformer with streaming Thinker–Talker architecture
License
Apache-2.0
Type
multimodal
Modalities
textimageaudiovideo

Benchmark Scores

Advanced Specifications

Model Family
Gander
Finetuned From
MiniCPM-o 4.5
API Access
Not Available
Chat Interface
Not Available
Variants
Gander-Unit8released Thinker and Talker pair

Capabilities & Limitations

Capabilities
full-duplex conversationstreaming speechaudio-visual understandinginterruption handlingtask delegation
Known Limitations
Long-horizon task execution depends on a configured external reasoning agentStreaming temporal units are not a standard token context limitPerception and factual errors remain possible
Notable Use Cases
realtime voice assistantsscreen-grounded conversational agentsmultimodal interaction research
Tool Use Support
Yes

Related Models