Chatterbox
Chatterbox is an open-source text-to-speech model from Resemble AI under the MIT license. In blind tests most listeners preferred it over ElevenLabs; it offers emotion control, voice cloning and low latency.
Modelle mit offen verfügbaren Gewichten — frei herunterladbar, selbst hostbar und feinabstimmbar, mit voller Kontrolle über Daten und Betrieb (Lizenzen variieren). Anbieterübergreifende Eigenschaft, z. B. FLUX.2, Llama und DeepSeek.
Chatterbox is an open-source text-to-speech model from Resemble AI under the MIT license. In blind tests most listeners preferred it over ElevenLabs; it offers emotion control, voice cloning and low latency.
Fish Audio S2 is an open-source text-to-speech model from the Fish Speech family. With inline tags it controls emotion and emphasis at word level, supports more than 80 languages and delivers very low latency.
FLUX.2 is the text-to-image model family from Black Forest Labs. It spans the proprietary Pro and Flex variants plus the open 32-billion-parameter Dev model and the compact Klein series, unifying image generation and image editing in a single model.
Foundation-Sec-8B is Cisco Foundation AIs open, cybersecurity-specialized language model (2025). It extends Llama-3.1-8B via continued pretraining on about 5 billion security tokens and is available under the permissive Apache-2.0 license.
Hunyuan3D 2.0 is Tencents open generative 3D model (January 2025). In two stages it produces high-resolution, textured 3D assets from text or image and supports text-, image-, sketch- and portrait-to-3D. The weights are freely available on Hugging Face.
Kokoro is an open text-to-speech model (v1.0, January 2025) with just 82M parameters. It produces natural-sounding speech in 8 languages with 54 voices, runs on around 1 GB of VRAM and ships under the Apache-2.0 license.
Microsoft TRELLIS.2 is the strongest open image-to-3D model (4B parameters, MIT license). It generates high-resolution assets with native PBR materials in seconds and supports Gaussian Splatting; weights are freely available on Hugging Face.
Canary is NVIDIAs open ASR and speech-translation model family (CC-BY-4.0) that tops the Open ASR Leaderboard on English accuracy ahead of OpenAI Whisper and covers 25 European languages.
SauerkrautLM is a German-language LLM family from the German startup VAGOsolutions — fine-tunes based on Llama, Qwen, Mistral, and other open architectures.