Command A Vision
Command A Vision adds enterprise image understanding to a Command A language backbone through a multimodal adapter and SigLIP2 vision encoder. The 112B-parameter model accepts text and images and generates text for OCR, charts, tables, document questions, and visual analysis. It officially targets six languages: English, Portuguese, Italian, French, German, and Spanish. The hosted API documents a 128K context and 8K output; the downloadable Hugging Face configuration defaults to 32K while supporting extension to 128K. Its downloadable weights use non-commercial terms. Cohere’s limitations explicitly exclude tool use despite the capability badges on the documentation page.
2025-07-31
112B
Dense autoregressive Transformer with SigLIP2 vision encoder and multimodal adapter
CC-BY-NC-4.0 with acceptable-use addendum
Specifications
- Parameters
- 112B
- Architecture
- Dense autoregressive Transformer with SigLIP2 vision encoder and multimodal adapter
- License
- CC-BY-NC-4.0 with acceptable-use addendum
- Context Window
- 128,000 tokens
- Max Output
- 8,000 tokens
- Training Data Cutoff
- June 1, 2024
- Type
- multimodal
- Modalities
- textimage
Benchmark Scores
Advanced Specifications
- Model Family
- Command
- API Access
- Available
- Chat Interface
- Available
- Multilingual Support
- Yes
- Variants
- command-a-vision-07-2025
Capabilities & Limitations
- Capabilities
- visionOCRdocument understandingtable understandingmultilingual
- Known Limitations
- Downloaded weights are licensed for non-commercial use; commercial deployments require separate terms.Tool use is explicitly unsupported by the official limitations.Downloadable configuration defaults to 32K; 128K requires configuration changes.Accepts images but does not generate them.
- Notable Use Cases
- enterprise document analysisvisual question answeringchart interpretation
- Function Calling Support
- No
- Tool Use Support
- No