Cohere logo

Aya Vision

CohereOpen WeightsPending Human Review

Aya Vision is Cohere’s multilingual vision-language family with 8B and 32B language tiers. It combines an autoregressive language model, a SigLIP2 vision encoder, and a multimodal adapter to handle image-and-text prompts in 23 languages. Its capabilities include OCR, image captioning, visual questions, translation of text in images, and document understanding. Both variants support a 16K context and generate text rather than images. Research weights use non-commercial terms. Cohere’s current Chat API serves the 32B checkpoint with a 4K output limit; the 8B hosted checkpoint was retired in April 2026.

2025-03-04
8B and 32B language tiers
Autoregressive Transformer with SigLIP2 vision encoder and multimodal adapter
CC-BY-NC-4.0 with acceptable-use addendum

Specifications

Parameters
8B and 32B language tiers
Architecture
Autoregressive Transformer with SigLIP2 vision encoder and multimodal adapter
License
CC-BY-NC-4.0 with acceptable-use addendum
Context Window
16,000 tokens
Max Output
4,000 tokens
Type
multimodal
Modalities
textimage

Benchmark Scores

Advanced Specifications

Model Family
Aya
API Access
Available
Chat Interface
Available
Multilingual Support
Yes
Variants
aya-vision-8b: 16K context; weights availableaya-vision-32b: 16K context; current hosted API

Capabilities & Limitations

Capabilities
vision23-language supportOCRvisual question answeringimage-text translation
Known Limitations
Downloaded weights are licensed for non-commercial use; commercial deployments require separate terms.Accepts images but does not generate them.The 8B hosted endpoint was retired April 4, 2026; weights remain available.Visual interpretations and translated image text can be incorrect.
Notable Use Cases
multilingual visual assistantsdocument understandingimage captioning research

Related Models