KI-Modelle
Konkrete KI-Modelle und Modellfamilien im Überblick — wer sie entwickelt, wie sie sich in Reasoning, Coding, Multimodalität und Preis unterscheiden und wofür die einzelnen Linien im Alltag am besten taugen. Gegliedert nach Anbieter-Familien (GPT, Claude, Gemini …) und übergreifenden Einsatzzwecken wie Security, Bild & Video oder Open-Weight.
- $/MTok (Cost per Million Tokens) API-Nutzung LLM-Pricing
Standard pricing unit for AI APIs — cost in US dollars per one million processed tokens. Listed separately for input, output and sometimes cache.
- AssemblyAI Universal-2 Speech & Voice
Universal-2 is AssemblyAIs closed speech-to-text model with industry-leading accuracy (mean WER around 6 percent) and strong formatting, punctuation and proper-noun recognition across 99 languages.
- Batch API API-Nutzung LLM-Pricing
Asynchronous API mode that collects many requests and processes them at a significant discount — results are typically delivered within 24 hours.
- Chatterbox Speech & Voice Open-Weight
Chatterbox is an open-source text-to-speech model from Resemble AI under the MIT license. In blind tests most listeners preferred it over ElevenLabs; it offers emotion control, voice cloning and low latency.
- Claude Fable 5 Claude
Claude Fable 5 is Anthropics first publicly available Mythos-class model, released in June 2026. It is built for the most demanding reasoning and long-horizon agentic work and offers a one-million-token context window.
- Claude Fable 5.1 Claude
Anthropic's flagship Claude Fable model, released on 1 September 2026. A frontier model for demanding reasoning, coding, research and long-running agents; 1M-token context, multimodal input, proprietary access via API and the Claude apps.
- Claude Haiku 4.5 Claude
Claude Haiku 4.5 is Anthropics fast compact model from October 2025 — it delivers near-Sonnet-4 performance at roughly one third of the cost and is the first Haiku model with reasoning.
- Claude Mythos Claude
Claude Mythos (April 2026) is Anthropic's most powerful model to date — a new tier above the Opus line, distributed primarily through the Glasswing program to cybersecurity defenders.
- Claude Mythos 5 Claude
Claude Mythos 5 (June 9, 2026) is Anthropic's most powerful model line above Opus — technically identical to Fable 5 but without its safety classifiers, distributed through Project Glasswing.
- Claude Mythos 5.1 Claude
Claude Mythos 5.1 (September 2026) is Anthropics frontier model for vetted cyber and life-science organizations — the same model as Fable 5.1, only with more permissive safeguards.
- Claude Opus 4.5 Claude
Claude Opus 4.5 is Anthropic's top-tier model from November 2025 — the direct predecessor of Opus 4.6, bringing "Infinite Chats" and substantially reduced token usage.
- Claude Opus 4.8 Claude
Claude Opus 4.8 is the flagship model of Anthropics Opus line, released in May 2026. It targets agentic coding, complex reasoning and computer use, and is described as an incremental but measurable improvement over Opus 4.7.
- Claude Opus 5 Claude
Anthropic's flagship Claude Opus model, released on 24 July 2026. Positioned as a powerful everyday model for knowledge work, coding and agents; proprietary access via the Claude apps and API (model ID claude-opus-5).
- Claude Sonnet 4.5 Claude
Claude Sonnet 4.5 (September 2025) is Anthropics Sonnet-line model tuned for agents and coding — leading on SWE-bench Verified at launch, since superseded by newer Sonnet versions.
- Claude Sonnet 4.6 Claude
Claude Sonnet 4.6 is Anthropics mid-tier model from February 2026 — close to Opus-level performance, with a 1M-token context and unchanged Sonnet pricing of 3/15 USD per million tokens.
- Claude Sonnet 5 Claude
Claude Sonnet 5 is Anthropic's mid-tier model from June 30, 2026 — the most agentic Sonnet yet, with performance close to Opus 4.8 at $2 input and $10 output per 1M tokens.
- Claude Sonnet 5.5 Claude
Claude Sonnet 5.5 (September 28, 2026) is Anthropic's low-cost mid-tier model — a faster, cheaper companion to Opus 5.5, priced at $2/$10 per 1M tokens and leading on Terminal-Bench 4.0. Available via boostN.
- Code Llama Llama
Code Llama is Metas coding-specialized Llama 2 variant from August 2023 — open-weights, in several sizes and with base, Python and Instruct flavors.
- Codestral Mistral
Codestral is Mistral AI's code-specialized model (first released May 2024) — a 22B-parameter model for code completion and generation, originally published as open weights.
- Deepgram Nova-3 Speech & Voice
Nova-3 is Deepgrams real-time speech-to-text model (released February 2025) with very low streaming latency of around 200 to 300 milliseconds, keyterm prompting and support for more than 30 languages.
- DeepSeek Janus-Pro-7B DeepSeek
Janus-Pro-7B is DeepSeek's open-source multimodal model from January 2025 — unifying image understanding and image generation in a single 7B architecture, MIT-licensed.
- DeepSeek R1 DeepSeek
Open reasoning model from the Chinese lab DeepSeek, released on 20 January 2025 under the MIT license. A mixture-of-experts with 671B parameters (37B active per token), heavily trained via reinforcement learning; the first open model to match OpenAI's o1 level.
- DeepSeek V3 DeepSeek
DeepSeek V3 is an open model from the Chinese provider DeepSeek, released in December 2024. It uses a Mixture-of-Experts architecture with 671 billion parameters (37 billion active per token) and a context window of around 128,000 tokens.
- DeepSeek V3.1 DeepSeek
DeepSeek V3.1 is DeepSeeks open-weights hybrid model from August 2025 — it unifies a think and a non-think mode in a single model, with a 128K context and strong agentic skills.
- DeepSeek V3.2 DeepSeek
DeepSeek V3.2 (December 1, 2025) is an open MoE model with 685B parameters under an MIT license — especially efficient on long contexts thanks to DeepSeek Sparse Attention.
- DeepSeek V4 DeepSeek
DeepSeek V4 (April 2026) is the fourth generation of the Chinese open-weight MoE model — released as Pro (1.6T parameters) and Flash (284B) under MIT license, with a 1M-token context window.
- DeepSeek V4-Flash DeepSeek
DeepSeek V4-Flash (July 2026) is an open MoE model with 284B parameters (13B active), a 1M-token context and an MIT license — freely downloadable on Hugging Face.
- DeepSeek V4-Pro DeepSeek
DeepSeek V4-Pro (2026) is DeepSeeks open-weight flagship — a 1.6-trillion-parameter MoE with about 49B active parameters per token, a 1M-token context and an MIT license.
- DeepSeek V4.1-Flash DeepSeek
DeepSeek V4.1-Flash (September 10, 2026) is the smallest model in the new DeepSeek V4.1 architecture family — the first DeepSeek generation with native multimodal image processing. A speed-/cost-optimized "Flash" model with open weights (MIT), available via boostN.
- Devstral Mistral
Devstral is Mistral AI's open agentic model for software engineering (May 2025) — trained to solve real GitHub issues and strong on SWE-bench Verified.
- ElevenLabs v3 Speech & Voice
ElevenLabs v3 is a highly expressive text-to-speech model with inline audio tags for emotion, emphasis and non-verbal sounds. It supports more than 70 languages and multi-speaker dialogue and is seen as a benchmark for realism.
- Fish Audio S2 Speech & Voice Open-Weight
Fish Audio S2 is an open-source text-to-speech model from the Fish Speech family. With inline tags it controls emotion and emphasis at word level, supports more than 80 languages and delivers very low latency.
- FLUX.2 Bild & Video Open-Weight
FLUX.2 is the text-to-image model family from Black Forest Labs. It spans the proprietary Pro and Flex variants plus the open 32-billion-parameter Dev model and the compact Klein series, unifying image generation and image editing in a single model.
- Foundation-Sec-8B Security-Modelle Open-Weight
Foundation-Sec-8B is Cisco Foundation AIs open, cybersecurity-specialized language model (2025). It extends Llama-3.1-8B via continued pretraining on about 5 billion security tokens and is available under the permissive Apache-2.0 license.
- Gemini 3 Flash Gemini
Gemini 3 Flash is Google's fast, low-cost thinking model from December 17, 2025 — near-Pro reasoning with a 1M-token context, built for agentic workflows and coding.
- Gemini 3 Pro Gemini
Gemini 3 Pro (November 2025) was Google's third flagship model with a 1M-token context window and native multimodality — now superseded by Gemini 3.1 Pro.
- Gemini 3.1 Pro Gemini
Gemini 3.1 Pro (February 2026) is Google's current flagship — 1M-token context, 64K output tokens, doubled reasoning performance over Gemini 3 Pro.
- Gemini 3.5 Flash Gemini
Gemini 3.5 Flash is Googles high-efficiency multimodal model from May 2026 — it brings near-Pro coding and reasoning quality at Flash-typical cost and speed.
- Gemini 3.6 Flash Gemini
Gemini 3.6 Flash (July 2026) is Googles fast, cost-efficient Flash tier with a 1M-token context — built for high throughput and low cost per request.
- Gemini 3.7 Flash Gemini
Gemini 3.7 Flash (August 2026) is Googles fast workhorse model for coding and agents — 1M-token context, introductory price of 0.75/3.75 USD per million tokens, half the cost of 3.6 Flash.
- Gemini 3.8 Flash Gemini
Gemini 3.8 Flash (September 2026) is Googles most intelligent workhorse model for coding, agents and multi-step reasoning — 1M-token context, introductory price of 0.75/3.75 USD per 1M tokens, direct successor to 3.7 Flash.
- Gemma 3 Gemma
Gemma 3 is Googles family of open models released in March 2025. It comes in four sizes (1B, 4B, 12B, 27B), is multimodal from 4B upward, supports over 140 languages and offers a context window of up to 128,000 tokens.
- Gemma 3n Gemma
Google's open, mobile-first AI model, available in preview from May 2025 and fully from 26 June 2025. Multimodal (text, image, audio, video), optimised for on-device use; open weights in the E2B and E4B sizes.
- Gemma 4 Gemma
Gemma 4 is Google DeepMind's open model family from April 2026 — four sizes (E2B to 31B), 256k context, multimodal, Apache 2.0 licensed.
- GLM-4.5 Zhipu AI
GLM-4.5 is an open model from Zhipu AI (Z.ai) released in July 2025. It is an agent-native Mixture-of-Experts model with 355 billion parameters (32 billion active), a 128,000-token context and an MIT licence.
- GLM-4.6 Zhipu AI
GLM-4.6 is Zhipu AI's open-weight MoE model from September 2025 — 355B parameters (32B active), 200k context, MIT-licensed, with a focus on coding performance.
- GLM-4.7 Zhipu AI
GLM-4.7 (December 2025) is Zhipu AIs open-weight coding flagship — an MoE model with roughly 400B parameters, a 200K input context and the "Preserved Thinking" feature.
- GLM-5 Zhipu AI
Open model from Zhipu AI (Z.AI), released in February 2026 under the MIT license. A mixture-of-experts with around 744B parameters (about 40B active per token), 200K context, focused on coding and agents. Freely downloadable via Hugging Face.
- GLM-5.1 Zhipu AI
GLM-5.1 is Z.ai's open-source flagship from April 2026 — 744B parameters (40B active), 200k context, designed for agentic coding with up to 8 hours of autonomous runtime.
- GLM-5.2 Zhipu AI
GLM-5.2 (June 2026) is Zhipu AIs (Z.ai) open-weight model — an MoE with around 744B parameters (~40B active), a 1M-token context and MIT license, strong at agentic coding.
- GLM-5.3 Zhipu AI
GLM-5.3 is Zhipu AIs open-weights model from August 2026 — a MoE model focused on coding and cybersecurity that tops the open-model field on agentic benchmarks.
- Google Imagen Bild & Video
Google Imagen is the text-to-image model family from Google DeepMind. The current generation, Imagen 4, launched in 2025 in Fast, Generate and Ultra tiers, is available via the Gemini API and Vertex AI, and is known for strong typography and prompt adherence.
- GPT Astra GPT Security-Modelle
GPT-6 Astra is OpenAIs flagship from September 2026 — the first model at the Critical Preparedness tier, focused on computer use and cybersecurity.
- GPT Image 2 Bild & Video
GPT Image 2 is OpenAIs native image model (April 2026) that reasons before drawing, renders text very reliably and produces high-resolution photorealistic images.
- GPT Luna GPT
GPT Luna (GPT-5.6 Luna) is the smallest, fastest and most affordable model of OpenAIs GPT-5.6 family, released in July 2026. It is built for cost-sensitive, high-volume workloads and offers a context window of roughly 1.05 million tokens.
- GPT Sol GPT
GPT-5.6 Sol is OpenAIs flagship from July 2026 — the top model of a new solar-system family (Sol, Terra, Luna) focused on coding, science and cybersecurity.
- GPT Terra GPT
GPT Terra is the middle tier of OpenAIs GPT-5.6 family (July 2026), sitting between the efficient Luna and the flagship Sol — built for a strong price-performance balance with a 1M-token context.
- GPT-4o GPT
GPT-4o (May 2024) was OpenAI's first natively multimodal model — text, audio and image in a single model, with a real-time voice mode and substantially lower cost than GPT-4 Turbo.
- GPT-4o mini GPT
GPT-4o mini (July 2024) is OpenAI's small, low-cost variant of GPT-4o — positioned as the successor to GPT-3.5 Turbo, clearly stronger at a much lower price per token.
- GPT-5.2 GPT
GPT-5.2 is OpenAI's model generation from December 2025 — three variants (instant, thinking, Pro), a 400k context window and new state-of-the-art on professional knowledge work.
- GPT-5.3 GPT
GPT-5.3 is an iteration of OpenAI's GPT-5 line — the Instant variant launched on March 3, 2026 with a focus on factual accuracy, shorter preambles and a more natural conversational flow.
- GPT-5.3-Codex GPT
GPT-5.3-Codex is OpenAI's specialized agentic coding model (February 2026), the successor to GPT-5.2-Codex — 25% faster, with substantial gains on Terminal-Bench and OSWorld.
- GPT-5.3-Codex-Spark GPT
GPT-5.3-Codex-Spark is OpenAI's coding model from February 2026 — a smaller, real-time-optimized variant of GPT-5.3-Codex delivering over 1000 tokens per second on Cerebras hardware.
- GPT-5.4 GPT
GPT-5.4 is OpenAI's reasoning model from March 5, 2026 — the first mainline model with the coding capabilities of GPT-5.3-codex, native computer-use controls and 1M-token context.
- GPT-5.5 GPT
GPT-5.5 is OpenAI's flagship from April 23, 2026 — positioned as its "smartest and most intuitive model" and pitched as a step toward AI that works across tools on its own.
- GPT-6.1 Sol GPT
GPT-6.1 Sol (September 29, 2026) is OpenAIs low-cost coding and agentic model — the successor to GPT-6 Sol after just seven days, with near GPT-6 Astra intelligence at roughly a fifth of the price. Available via boostN.
- GPT-Live-1 GPT
OpenAI's real-time voice model, the default voice experience in ChatGPT since 8 July 2026. It listens and speaks at the same time (bidirectional), allows interruptions, and hands complex questions off to a stronger text model. Proprietary access via ChatGPT.
- Grok 3 Grok
Reasoning model from xAI (Elon Musk's AI company), released on 17 February 2025. Trained on the Colossus supercluster, with 'Think' reasoning and DeepSearch plus real-time access to X. Proprietary access via the Grok app and API.
- Grok 4 Grok
Grok 4 is xAI's reasoning model from July 2025 — text and image input, 256k context, trained via reinforcement learning on the 200,000-GPU Colossus cluster.
- Grok 4.1 Grok
Grok 4.1 is xAIs AI model released in November 2025. It delivers more natural dialogue, fewer hallucinations and better style control, and comes in two variants: one that answers directly and one that reasons before responding (Thinking).
- Grok 4.3 Grok
Grok 4.3 is xAIs reasoning-oriented model from April 2026 — with a 1M-token context, native video input and a focus on agentic workflows and factual accuracy.
- Grok 4.5 Grok
Grok 4.5 is xAI's July 2026 frontier model for coding, agentic tasks and knowledge work — high token efficiency at $2 input and $6 output per 1M tokens.
- Grok 4.6 Grok
Grok 4.6 (August 2026) is xAIs flagship model for long-running agentic work — 500k-token context, text and image input, tiered pricing from 2/6 USD per million tokens.
- Hunyuan3D 2.0 3D-Modelle Open-Weight
Hunyuan3D 2.0 is Tencents open generative 3D model (January 2025). In two stages it produces high-resolution, textured 3D assets from text or image and supports text-, image-, sketch- and portrait-to-3D. The weights are freely available on Hugging Face.
- Hyper3D Rodin 3D-Modelle
Hyper3D Rodin is a commercial 3D generation model by Deemos. Generation 2.5 turns images or text into highly detailed geometry with over 10 million polygons and 3D-native PBR textures, and is regarded as a leader in photorealism.
- Input Token API-Nutzung LLM-Pricing
Tokens you send to an AI model in an API call — your prompt, the context, attached documents. Billed separately from output tokens and usually much cheaper.
- Kimi K2 Moonshot AI
Kimi K2 is Moonshot AI's open-weight MoE model series with 1T total parameters and 32B active parameters, trained on 15.5T tokens — focused on agentic and coding tasks.
- Kimi K2 Thinking Moonshot AI
Kimi K2 Thinking (November 2025) is Moonshot AI's open reasoning-agent model — a trillion-parameter MoE with a 256K context that runs long thinking and tool-use chains on its own.
- Kimi K2.5 Moonshot AI
Kimi K2.5 is an open model from Moonshot AI released in January 2026. It is a native multimodal Mixture-of-Experts model with around 1 trillion parameters (32 billion active), a 256,000-token context and agentic capabilities.
- Kimi K2.6 Moonshot AI
Kimi K2.6 is Moonshot AI's open-weight agent model from April 2026 — 1T parameters (32B active), 262k context, natively multimodal, with agent swarms of up to 300 sub-agents.
- Kimi K3 Moonshot AI
Kimi K3 (July 2026) is Moonshot AIs open-weight flagship — an MoE model with 2.8T parameters (104B active), native vision and a 1M-token context.
- Kling 3.0 Bild & Video
Kling 3.0 is Kuaishous video model, released on February 5, 2026. It generates photorealistic clips of up to 15 seconds with native audio across multiple languages and is regarded as the strongest price-to-performance video generator.
- Kokoro Speech & Voice Open-Weight
Kokoro is an open text-to-speech model (v1.0, January 2025) with just 82M parameters. It produces natural-sounding speech in 8 languages with 54 voices, runs on around 1 GB of VRAM and ships under the Apache-2.0 license.
- Llama 2 Llama
Llama 2 is Metas open-weights language model from July 2023 — the first Llama with a commercially usable license, in 7B, 13B and 70B sizes with a 4K context.
- Llama 3 Llama
Meta's open model generation, released on 18 April 2024. It launched in the 8B and 70B sizes (each in base and Instruct variants) under the Llama 3 community license; freely downloadable, for research and commercial use.
- Llama 3.1 Llama
Llama 3.1 (July 2024) is Metas open-weight family in 8B, 70B and 405B with a 128K context and multilinguality — the 405B was the first open model on par with GPT-4o and Claude 3.5 Sonnet.
- Llama 3.2 Llama
Llama 3.2 is Meta's model generation from September 2024 — the first with vision support (11B/90B) and lightweight edge models (1B/3B) for mobile and on-device use.
- Llama 3.3 Llama
Llama 3.3 (December 2024) is Metas open 70B instruct model — it reaches 405B-class quality at 70B cost, with a 128K context and open weights.
- Llama 4 Maverick Llama
Llama 4 Maverick (April 2025) is Meta's large Mixture-of-Experts variant in the Llama 4 lineup — 17B active out of 400B parameters across 128 experts, 1M-token context.
- Llama 4 Scout Llama
Llama 4 Scout (April 2025) is Meta's smaller Mixture-of-Experts variant in the Llama 4 lineup — 17B active out of 109B parameters, 10M-token context, multimodal.
- Llama-3-SauerkrautLM-70b Llama
Llama-3-SauerkrautLM-70b is a German-language DPO fine-tune of Meta Llama 3 70B, developed by VAGOsolutions and Hyperspace.ai.
- Magistral Mistral
Magistral is Mistral AIs first reasoning model from June 2025 — with a transparent, verifiable chain of thought and strength in European languages, partly as open weights.
- Meshy 3D-Modelle
Meshy is a commercial hosted platform for 3D generation from text or images. Version 6 produces production-ready assets in about a minute with PBR textures, auto-retopology and rigging, and is regarded as a versatile all-round tool.
- Microsoft TRELLIS.2 3D-Modelle Open-Weight
Microsoft TRELLIS.2 is the strongest open image-to-3D model (4B parameters, MIT license). It generates high-resolution assets with native PBR materials in seconds and supports Gaussian Splatting; weights are freely available on Hugging Face.
- Midjourney v8 Bild & Video
Midjourney v8 is the eighth generation of Midjourneys image generator (alpha March 2026) with roughly five times faster generation, native 2K resolution and improved text rendering.
- MiniMax MiniMax
MiniMax is an AI lab founded in Shanghai in 2021 whose open-weight M model family (LLMs) targets very long context windows and low running costs via sparse attention.
- Mistral 7B Mistral
Mistral 7B (September 2023) is Mistral AIs first open-weight model — a 7.3B-parameter model under Apache 2.0 that at the time outperformed the larger Llama 2 13B on benchmarks.
- Mistral Large 3 Mistral
Mistral Large 3 (December 2025) is Mistral's open-weight MoE flagship — 675B parameters, 256K-token context, natively multimodal, Apache 2.0.
- Mistral Medium 3 Mistral
Mistral Medium 3 is a proprietary AI model from Mistral AI released in May 2025. As an enterprise model built for high performance at low operating cost, it offers around 128,000 tokens of context and can also be run on-premise or in your own VPC.
- Mistral Small 4 Mistral
Mistral Small 4 is Mistral AI's open-weight MoE model from March 2026 — merging reasoning, vision and coding into one model, 119B parameters (6.5B active), Apache 2.0.
- Mixtral 8x7B Mistral
Mixtral 8x7B (December 2023) is Mistral AIs open Sparse-MoE model with 8 experts — 46.7B parameters, 12.9B active per token, 32K context, Apache 2.0 license.
- Muse Spark Muse
Muse Spark is Metas closed-weight flagship from April 2026 — the first model of the Muse line out of Meta Superintelligence Labs, focused on efficient reasoning via thought compression.
- NVIDIA Canary Speech & Voice Open-Weight
Canary is NVIDIAs open ASR and speech-translation model family (CC-BY-4.0) that tops the Open ASR Leaderboard on English accuracy ahead of OpenAI Whisper and covers 25 European languages.
- Opus 4.6 Claude
Claude Opus 4.6 is Anthropic's flagship model from February 2026 — top tier of the Claude 4 family with a 1M-token context window, stronger coding and new effort controls.
- Opus 4.7 Claude
Claude Opus 4.7 is Anthropic's top-tier model from April 2026 — successor to Opus 4.6 with high-resolution image processing, a new xhigh effort level and task budgets for long agentic runs.
- Output Token API-Nutzung LLM-Pricing
Tokens an AI model produces as its response. Billed separately and usually three to five times more expensive than input tokens because the model has to actively generate them.
- Pixtral Mistral
Mistral AI's first multimodal model, released in September 2024 under Apache 2.0. A vision-language model with 12B parameters plus a 400M vision encoder, processing text and (multiple) images; 128K context, open weights via Hugging Face.
- Prompt Caching API-Nutzung LLM-Pricing
Prompt caching is an API feature in which a provider stores recurring prompt prefixes — making subsequent requests cheaper and faster because the cached portion is not reprocessed.
- Qwen 2.5 Qwen
Open model family from Alibaba, released in September 2024. It spans base sizes from 0.5B to 72B parameters plus Coder and Math variants; 128K context, multilingual. Smaller sizes under Apache 2.0, the 72B model under a dedicated Qwen license.
- Qwen 3 Qwen
Qwen 3 is Alibabas family of open models released in April 2025. It ranges from 0.6B to 235B parameters (dense and MoE models), is licensed under Apache 2.0 and unites a thinking and a non-thinking mode in a single model.
- Qwen 3.5 Qwen
Qwen 3.5 (February 2026) is Alibaba's third-generation open model family — from 0.8B to 397B parameters, MoE flagship with 1M-token context, multimodal.
- Qwen 3.6 Qwen
Qwen 3.6 is Alibabas open-weights model from April 2026 — the dense 27B variant matches much larger MoE models on agentic coding benchmarks, under an Apache 2.0 license.
- Qwen 3.7-Max Qwen
Qwen 3.7-Max (May 2026) is Alibabas proprietary Max flagship with a 1M-token context, positioned as the "Agent Frontier" for long, tool-intensive agent workflows.
- Qwen 3.8-Max Qwen
Qwen 3.8-Max (August 2026) is Alibabas largest flagship — an MoE model with 2.4T parameters, a 1M-token context and multimodal input, later released with open weights.
- Qwen3-Max Qwen
Qwen3-Max is Alibaba's trillion-parameter flagship from September 2025 — a Mixture-of-Experts model with a 262K context that, unlike earlier Qwen models, is API-only.
- Rate Limit (AI) API-Nutzung LLM-Pricing
Provider-enforced cap on requests or tokens per time window — it protects infrastructure and ensures fair usage across customers.
- Runway Gen-4.5 Bild & Video
Runway Gen-4.5 is Runways video model that turns text or an image into 5- and 10-second clips with strong prompt adherence and cinematic motion. It was built with NVIDIA and uses an Autoregressive-to-Diffusion technique.
- SauerkrautLM Open-Weight
SauerkrautLM is a German-language LLM family from the German startup VAGOsolutions — fine-tunes based on Llama, Qwen, Mistral, and other open architectures.
- Sec-Gemini v1 Security-Modelle
Sec-Gemini v1 is Googles experimental, Gemini-based AI model for cybersecurity. It combines Gemini reasoning with current threat intelligence from Mandiant, the OSV vulnerability database and Google Threat Intelligence.
- Seedream 4.0 Bild & Video
Seedream 4.0 is ByteDances multimodal image model (September 2025) that processes text and multiple images as input, renders up to 4K and runs more than ten times faster than Seedream 3.0.
- Tier Pricing API-Nutzung LLM-Pricing
Tiered model line-up from a provider — small fast variants (Mini/Flash/Haiku) at a fraction of the price of the big frontier models. Also: volume tiers with quantity discounts.
- Tripo AI 3D-Modelle
Tripo AI is a commercial 3D generation platform by VAST AI. Version 3.1 leads the 2026 ELO arena rankings and turns text or images into game-ready models with clean quad topology, PBR textures and auto-rigging.
- Veo 3.1 Bild & Video
Veo 3.1 is Googles video model (DeepMind) that turns text or an image into 8-second clips with natively synchronized audio. Since the January 2026 update it delivers true 4K (3840x2160) and native vertical formats.
- Whisper LLM-Grundlagen Speech & Voice
Whisper is OpenAI's open speech-to-text model from 2022 — a multilingual encoder-decoder transformer that turns audio into text and ships in several sizes under an MIT license.