AI in Review — August 2026: A Price U-Turn, the Open-Weight Wave and an Enforceable EU AI Act
August 2026 in AI: Sonnet 5 price freeze, Qwen3.8-Max, the open-weight wave, Google AI Mode on Gemini Flash and the enforceable EU AI Act — sorted out.
in AI Models
What do AI models cost? Token prices by input vs. output, surcharges for tool use and reasoning, batch discounts, and savings from prompt caching — plus how to realistically estimate the cost per request.
August 2026 in AI: Sonnet 5 price freeze, Qwen3.8-Max, the open-weight wave, Google AI Mode on Gemini Flash and the enforceable EU AI Act — sorted out.
Real numbers from a customer setup with Claude Sonnet. What prompt caching delivers, when it pays off — and where the pitfalls are.
AI providers are pushing flat rates toward metered billing. Why it had to happen, what it costs and three levers that soften the shift.
Standard pricing unit for AI APIs — cost in US dollars per one million processed tokens. Listed separately for input, output and sometimes cache.
Tokens you send to an AI model in an API call — your prompt, the context, attached documents. Billed separately from output tokens and usually much cheaper.
Tokens an AI model produces as its response. Billed separately and usually three to five times more expensive than input tokens because the model has to actively generate them.
Asynchronous API mode that collects many requests and processes them at a significant discount — results are typically delivered within 24 hours.
Prompt caching is an API feature in which a provider stores recurring prompt prefixes — making subsequent requests cheaper and faster because the cached portion is not reprocessed.
Provider-enforced cap on requests or tokens per time window — it protects infrastructure and ensures fair usage across customers.
Tiered model line-up from a provider — small fast variants (Mini/Flash/Haiku) at a fraction of the price of the big frontier models. Also: volume tiers with quantity discounts.
How AI models bill — tokens, input vs. output, hidden cost drivers and three levers to save. With price table and worked examples.
Anthropic released Claude Fable 5.1 and Mythos 5.1. Base price holds, but cache reads are 75% cheaper and token usage drops by half.
Anthropic released Claude Opus 5.5: $4 / $20 per million tokens, about 40% cheaper to run than Opus 5 and ahead of Fable 5.1 on coding benchmarks.
OpenAI released GPT-6.1 Sol: $2 / $10 per million tokens, near-Astra coding performance, but behind Anthropic's models on the overall index.
Anthropic has ended the 50% boost. Since September 14 at 9:00 CEST, the permanent baseline is +25% – roughly 17% less capacity than before.
Anthropic releases Claude Sonnet 5 on July 1, 2026 — near-Opus performance per the vendor, at an intro price of 2/10 dollars per million tokens.
The 50 percent hike on Claude Sonnet 5 set for September 1 is off. Anthropic makes the 2/10 dollar introductory price permanent.