Back to glossary

Term

Pixtral

Mistral AI's first multimodal model, released in September 2024 under Apache 2.0. A vision-language model with 12B parameters plus a 400M vision encoder, processing text and (multiple) images; 128K context, open weights via Hugging Face.

Pixtral — explained in more detail

Pixtral (Pixtral 12B) is Mistral AI’s first multimodal model, released in September 2024 under the open Apache 2.0 license. As a vision-language model, it processes text and images in a shared context — a step beyond the previously text-only Mistral models. The weights are openly available and can be freely downloaded, self-hosted and fine-tuned.

Technically, Pixtral builds on Mistral’s text model Nemo 12B and adds a dedicated vision encoder of around 400 million parameters; together this makes a model of around 12 billion parameters plus an image adapter. It supports a context window of up to 128,000 tokens and can process multiple images per input — passed as URLs or as base64-encoded data. Despite its comparatively compact size, performance on text-only benchmarks remains competitive.

Access is open: Pixtral is available via Hugging Face (version Pixtral-12B-2409) and through major cloud platforms.

Example / practical relevance

Pixtral suits tasks where image and text must be considered together: image captioning, object counting, reading charts or screenshots, and document understanding with figures. The ability to input several images at once is useful, for example, for comparisons or image series.

For teams with data-protection or cost requirements, the Apache 2.0 license is attractive: the model can run locally, with no per-request API cost and without image data going to an external service.

Distinction from similar terms

Pixtral is Mistral’s multimodal line and differs from the text-only Mistral models (such as Mistral Nemo or Mistral Large). Compared with proprietary vision models (GPT-4o, Gemini), Pixtral is openly licensed and self-hostable. The name combines “pixel” and “Mistral” and marks the image focus of the line.

See everything in one place:Pixtral