核心信息
腾讯正式发布 AuK,一个面向统一语音生成与编辑的开源基础模型。它支持自然语言指令加参考音频,通过单一接口完成各类语音任务。
要点
- 支持零样本 TTS、指令控制生成、内容编辑、语音转换、去口音、音色/风格/情感编辑,以及语速/音高控制。
- 同时支持增强、降噪、多说话人和音乐分离等功能。
- 同步推出 AuK-Flash:4 步推理,在同等条件下速度约快 4.5 倍;代码、权重等资源已发布。
腾讯正式发布 AuK,一个面向统一语音生成与编辑的开源基础模型。它支持自然语言指令加参考音频,通过单一接口完成各类语音任务。
🚀 AuK is officially here. Nano banana🍌 for audio An open-source foundation model for unified speech generation and editing. Natural-language instructions + reference audio. One interface. Zero-shot TTS. Instruction-controlled generation. Content editing. Whisper-conversion. De-accent. Timbre/style/emotion edit. Speed/Pitch control. Enhancement, denoising, multi-speaker and music separation. Also releasing AuK-Flash: 4-step inference. ~4.5× faster under matched conditions. Code, weights, and demo are live. Try it and share your feedback. 🤗 Paper & upvote: https://t.co/zEveUsuJRF ⭐ GitHub & star: https://t.co/72m9Msk8Yv