Cohere logo

Command A Vision

CohereOpen WeightsPending Human Review

Command A Vision adds enterprise image understanding to a Command A language backbone through a multimodal adapter and SigLIP2 vision encoder. The 112B-parameter model accepts text and images and generates text for OCR, charts, tables, document questions, and visual analysis. It officially targets six languages: English, Portuguese, Italian, French, German, and Spanish. The hosted API documents a 128K context and 8K output; the downloadable Hugging Face configuration defaults to 32K while supporting extension to 128K. Its downloadable weights use non-commercial terms. Cohere’s limitations explicitly exclude tool use despite the capability badges on the documentation page.

2025-07-31
112B
Dense autoregressive Transformer with SigLIP2 vision encoder and multimodal adapter
CC-BY-NC-4.0 with acceptable-use addendum

Specifications

Parameters
112B
Architecture
Dense autoregressive Transformer with SigLIP2 vision encoder and multimodal adapter
License
CC-BY-NC-4.0 with acceptable-use addendum
Context Window
128,000 tokens
Max Output
8,000 tokens
Training Data Cutoff
June 1, 2024
Type
multimodal
Modalities
textimage

Benchmark Scores

Advanced Specifications

Model Family
Command
API Access
Available
Chat Interface
Available
Multilingual Support
Yes
Variants
command-a-vision-07-2025

Capabilities & Limitations

Capabilities
visionOCRdocument understandingtable understandingmultilingual
Known Limitations
Downloaded weights are licensed for non-commercial use; commercial deployments require separate terms.Tool use is explicitly unsupported by the official limitations.Downloadable configuration defaults to 32K; 128K requires configuration changes.Accepts images but does not generate them.
Notable Use Cases
enterprise document analysisvisual question answeringchart interpretation
Function Calling Support
No
Tool Use Support
No

Related Models