跳到正文
RCreddit.com·

tencent/AuK-Flash · Hugging Face

AI 摘要

AuK-Flash is a distilled variant of the AuK Base model, offering fast 4-step inference for speech generation and editing. It supports various tasks including Zero-shot TTS, Instruct TTS, and content editing features like rewriting text, lyrics, and adjusting pitch, speed, and volume. Additionally, AuK-Flash provides paralinguistic editing for emotion and timbre, de-accenting, nonverbal editing, whisper conversion, and enhancement/separation capabilities for speech and music.

时间与来源

时间显示为 UTC

显示时区:UTC

本地时区尚不可用,暂时显示 UTC。

发布当时偏移:UTC+02026年9月12日 13:17 UTC

收录当时偏移:UTC+02026年9月12日 22:01 UTC

发布
2026年9月12日 13:17
收录
2026年9月12日 22:01
来源类型
开发者社区
档位
社区
信源状态
正常

档位是按信源手工设定的编辑判断,不是逐条打分。

AuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface. AuK has two variants:

Model Description Weight AuK Base model for high-quality generation 🤗 Hugging Face · 🤖 ModelScope AuK-Flash Distilled model for fast 4-step inference 🤗 Hugging Face · 🤖 ModelScope This repository contains the official weights for AuK-Flash, the distilled variant with fast 4-step inference.

AuK exposes every task through the same natural-language instruction interface. The table below groups the supported tasks by category, with a short description and a link to its section in the Cookbook , which provides instruction templates plus CLI and Python examples.

Category Task Description Cookbook Speech Generation Zero-shot TTS Speak the target text in the voice of the reference audio. Zero-shot TTS Instruct TTS Generate speech from a voice description alone — no reference audio. Instruct TTS Content Editing Speech Content Editing Rewrite what is said — replace, insert, or remove text. Speech Content Editing Lyric Editing Rewrite lyrics in a singing recording while preserving the melody and voice. Lyric Editing Acoustic Editing Pitch Editing Raise or lower the pitch by semitones. Pitch Editing Speed Editing Adjust the speaking rate; output length scales with the speed factor. Speed Editing Volume Editing Raise or lower the volume by decibels. Volume Editing Paralinguistic Editing Emotion Change the emotion while preserving content and voice. Emotion Timbre Change the timbre to a description while keeping the content unchanged. Timbre De-accent Remove a regional accent while preserving the speaker's voice and content. De-accent Nonverbal Editing Remove or add nonverbal sounds such as breaths, laughs, or coughs. Nonverbal Editing Whisper Conversion Convert between normal speech and whisper while preserving speaker and content. Whisper Conversion Enhancement & Separation Speech Enhancement Denoise, dereverberate, or restore natural, clear speech. Speech Enhancement Speech Separation Keep one speaker by talking order and remove the others. Speech Separation Music Separation Extract the singing voice from a mix, or keep all human voices. Music Separation Target Speaker Extraction Keep the target speaker identified by what they say . Target Speaker Extraction

来源·reddit.com