Command A Vision adds enterprise image understanding to a Command A language backbone through a multimodal adapter and SigLIP2 vision encoder. The 112B-parameter model accepts text and images and generates text for OCR, charts, tables, document questions, and visual analysis. It officially targets six languages: English, Portuguese, Italian, French, German, and Spanish. The hosted API documents a 128K context and 8K output; the downloadable Hugging Face configuration defaults to 32K while supporting extension to 128K. Its downloadable weights use non-commercial terms. Cohere’s limitations explicitly exclude tool use despite the capability badges on the documentation page.
Typemultimodal
Parameters112B