01
TTS and speech-content editing
Zero-shot TTS uses reference audio plus target text, while instruction TTS can generate from a voice description without a reference clip. Voice cloning should only use reference voices for which the operator has appropriate consent and rights.
Speech-content editing changes what is said inside an existing recording. For creator work, verify timing, speaker consistency and artifacts before replacing the original take.
02
Acoustic and paralinguistic edits
Tencent documents pitch, speed and volume edits as well as emotion, timbre, de-accenting, nonverbal-sound changes and whisper conversion. Start with one transformation at a time so failures are easier to diagnose.
A documented task is not a promise of studio-quality output for every language, speaker or recording condition. Human listening review remains essential for consequential work.
03
Enhancement and separation
AuK supports denoising, dereverberation, speech separation, music separation and target-speaker extraction. That gives the same family roles often handled by dedicated restoration and source-separation systems.
Independent matched comparisons against specialist tools are still sparse, so do not assume AuK is best-in-class at every supported task simply because the interface is broad.
04
A safer creator workflow
Keep the original recording, make the smallest necessary edit, apply cleanup or paralinguistic transformations separately, compare against the source and retain an edit log for client or high-stakes work.
For setup use the local AuK guide; for node-based work use the ComfyUI guide; and return to the main overview for release and license details.
Sources
Primary and supporting sources
Facts were rechecked against the linked sources immediately before publication. Pricing, product availability and rollout status can change.