Google's new Flash TTS models let you design AI voices from scratch using text descriptions
AI Digest
Flash TTS模型语音克隆多语言支持场景化对话API计费
Google推出Flash TTS和Flash-Lite TTS模型,支持100+语言,用户可通过文本描述创建AI声音,兼具语音克隆和场景化对话功能,即将通过API开放。
Google's Flash TTS and Flash-Lite TTS models enable voice creation from text descriptions across 100+ languages, with voice cloning and scene-directed dialogue, set to launch via APIs.
Key points
- 支持100+语言的文本生成语音,可定制音色、口音等特征 Text-to-speech models support over 100 languages with customizable voice traits
- 提供语音克隆功能,30秒音频可生成个性化声纹 Voice cloning feature creates profiles from 30-second audio samples
- 支持场景化对话控制,包含笑声、叹气等非语言声音 Scene-directed dialogue controls include laughter and sighs
- 通过Gemini API和AI Studio逐步开放,企业版API即将推出 Models roll out via Gemini API with enterprise access pending
- 按需计费模式,音频输出成本约0.54-1.62美元/小时 Pricing based on text/audio tokens with cost estimates provided
Takeaway: AI语音生成进入高度定制化时代,文本描述即可创建独特声线。 / AI voice generation reaches new customization levels through text-based voice creation.
Why it matters 揭示AI语音生成技术突破,为开发者提供新工具和商业机会。
View original ↗ Back to hot list
This page is an aggregated digest from decoder; content and hot-score data come from public sources. Copyright belongs to the original authors. We link to originals with nofollow and never republish full text.