Skip to content
RCreddit.com·

tencent/AuK-Flash · Hugging Face

AI summary

AuK-Flash is a distilled variant of the AuK Base model, offering fast 4-step inference for speech generation and editing. It supports various tasks including Zero-shot TTS, Instruct TTS, and content editing features like rewriting text, lyrics, and adjusting pitch, speed, and volume. Additionally, AuK-Flash provides paralinguistic editing for emotion and timbre, de-accenting, nonverbal editing, whisper conversion, and enhancement/separation capabilities for speech and music.

Time & source

Times shown in UTC

Display time zone: UTC

Local time zone unavailable; showing UTC.

PublishedOffset at this time: UTC+0Sep 12, 2026, 13:17 UTC

IngestedOffset at this time: UTC+0Sep 12, 2026, 22:01 UTC

Published
Sep 12, 2026, 13:17
Ingested
Sep 12, 2026, 22:01
Source type
Dev community
Tier
Community
Source status
Healthy

Tier is a per-source editorial setting, not a per-item score.

AuK is a 1.5B foundation model for speech generation and editing. Trained on millions of hours of diverse audio data, AuK supports zero-shot and instruction-based TTS, content and acoustic editing, paralinguistic editing, speech enhancement, and source separation through a unified natural-language instruction interface. AuK has two variants:

Model Description Weight AuK Base model for high-quality generation 🤗 Hugging Face · 🤖 ModelScope AuK-Flash Distilled model for fast 4-step inference 🤗 Hugging Face · 🤖 ModelScope This repository contains the official weights for AuK-Flash, the distilled variant with fast 4-step inference.

AuK exposes every task through the same natural-language instruction interface. The table below groups the supported tasks by category, with a short description and a link to its section in the Cookbook , which provides instruction templates plus CLI and Python examples.

Category Task Description Cookbook Speech Generation Zero-shot TTS Speak the target text in the voice of the reference audio. Zero-shot TTS Instruct TTS Generate speech from a voice description alone — no reference audio. Instruct TTS Content Editing Speech Content Editing Rewrite what is said — replace, insert, or remove text. Speech Content Editing Lyric Editing Rewrite lyrics in a singing recording while preserving the melody and voice. Lyric Editing Acoustic Editing Pitch Editing Raise or lower the pitch by semitones. Pitch Editing Speed Editing Adjust the speaking rate; output length scales with the speed factor. Speed Editing Volume Editing Raise or lower the volume by decibels. Volume Editing Paralinguistic Editing Emotion Change the emotion while preserving content and voice. Emotion Timbre Change the timbre to a description while keeping the content unchanged. Timbre De-accent Remove a regional accent while preserving the speaker's voice and content. De-accent Nonverbal Editing Remove or add nonverbal sounds such as breaths, laughs, or coughs. Nonverbal Editing Whisper Conversion Convert between normal speech and whisper while preserving speaker and content. Whisper Conversion Enhancement & Separation Speech Enhancement Denoise, dereverberate, or restore natural, clear speech. Speech Enhancement Speech Separation Keep one speaker by talking order and remove the others. Speech Separation Music Separation Extract the singing voice from a mix, or keep all human voices. Music Separation Target Speaker Extraction Keep the target speaker identified by what they say . Target Speaker Extraction

Source·reddit.com