Term
Mixtral 8x7B
Mixtral 8x7B (December 2023) is Mistral AIs open Sparse-MoE model with 8 experts — 46.7B parameters, 12.9B active per token, 32K context, Apache 2.0 license.
Mixtral 8x7B — explained in more detail
Mistral AI released Mixtral 8x7B on December 11, 2023 as an open-source model under the Apache 2.0 license. It is a Sparse Mixture-of-Experts (SMoE) model with eight expert networks: a learned router selects two of the eight experts per token at each layer and combines their outputs. The model has 46.7B total parameters, of which only about 12.9B are activated per forward pass — hence its high efficiency. The context window spans 32,768 tokens.
Example / Practical use
At release, Mixtral 8x7B outperformed Llama 2 70B on most benchmarks at roughly six times faster inference, and matched or beat GPT-3.5 on common standard benchmarks. Because only a fraction of the parameters is active per token, the model delivers the quality of a much larger model at the compute cost of a smaller one. As an Apache 2.0 model with open weights, it can be freely downloaded, adapted and self-hosted — an important building block of the early open-weight wave.
Distinction
Within the Mistral family, Mixtral stands for the MoE architecture — unlike the dense base model Mistral 7B, where all parameters are active. Compared with later, larger models (such as Mistral Large), Mixtral 8x7B is the early, efficiency-oriented MoE model. The name “8x7B” refers to eight experts based on the 7B architecture, but due to shared components it amounts to 46.7B total parameters rather than 56B.
Discover more
GPT-6 Astra Found Questions My First Security Review Missed
I used GPT-6 Astra and Fable 5.1 as independent reviewers for RLS, API and tenant-isolation checks. The useful part was the structured cross-review.
GlossaryCodestral
Codestral is Mistral AI's code-specialized model (first released May 2024) — a 22B-parameter model for code completion and generation, originally published as open weights.
EncyclopediaComparing AI models — who builds what and how to choose
The major model families in 2026 at a glance. Who builds Claude, GPT, Gemini, Llama, Mistral, DeepSeek, Qwen — and which model to pick when.