Back to glossary

Term

Llama 3.3

Llama 3.3 (December 2024) is Metas open 70B instruct model — it reaches 405B-class quality at 70B cost, with a 128K context and open weights.

Llama 3.3 — explained in more detail

Meta released Llama 3.3 70B Instruct on December 6, 2024 as the final model of its 2024 Llama release cadence. It is a text-only, instruction-tuned transformer model with 70.6B parameters that uses Grouped-Query Attention (GQA) for more scalable inference. The context window spans 128K tokens, the knowledge cutoff is December 2023, and it was trained on over 15T tokens of publicly available data. It supports eight languages, including English, German, French and Spanish.

Example / Practical use

The core advantage of Llama 3.3: it targets 405B-class results at the serving cost of a 70B model. According to Meta, Llama 3.3 70B generates responses nearly five times more cost-efficiently than the larger Llama 3.1 405B — with infrastructure costs of about 10 cents per 1M input tokens and 40 cents per 1M output tokens. As an open-weight model it is freely downloadable (Hugging Face) and available through platforms such as Amazon Bedrock; this makes it suitable for cost-sensitive deployments and self-hosting.

Distinction

Llama 3.3 is a text-only model and thus narrower than multimodal models. Within the Llama range it is the efficiency-optimized 70B variant that approaches the quality of the much larger 3.1-405B generation without its compute load. Unlike proprietary models from OpenAI or Anthropic, Llama is released under Metas open community license.

See everything in one place:Llama