
Run Qwen3-TTS Locally: Setup, Voice Cloning and Hardware
Run Qwen3-TTS locally with official Python tools or a desktop workflow. Compare checkpoints, Mac benchmark results, reference audio and setup pitfalls.
Qwen3-TTS has downloadable models for local speech generation. Choose Base for reference-audio cloning, CustomVoice for built-in speakers, or VoiceDesign for describing a new voice. Those are different checkpoints, not interchangeable modes of the same file. This guide covers the official setup and the smaller quantized model used by Local AI Audio.
Checked on October 5, 2026. Performance numbers below come from our September 23 test, not a new test of every current installer. Qwen-Audio-3.1 is a different model family: see cloud APIs versus local Qwen3-TTS.
Pick a route and checkpoint
| Goal | Model / route | What to check first |
|---|---|---|
| Read in a built-in voice | Official CustomVoice checkpoint | Supported speaker names and languages |
| Clone a reference voice | Official 0.6B or 1.7B Base checkpoint | Clean reference recording and matching transcript |
| Describe a new speaker | Official 1.7B VoiceDesign checkpoint | A separate model download |
| Use the desktop speech workflow | Local AI Audio's Qwen3-TTS 0.6B Base Q8 via audio.cpp | Model availability in your installed release |
The official repository documents the Python route. The 0.6B Base model card identifies the smaller cloning checkpoint and its Apache-2.0 license. A downloaded weight file is not the whole runtime: keep the tokenizer and dependencies required by your chosen backend.
Official local setup
Use a fresh environment with a CUDA-compatible PyTorch installation for the documented NVIDIA examples. The upstream quickstart recommends Python 3.12:
conda create -n qwen3-tts python=3.12 -y
conda activate qwen3-tts
python -m pip install -U qwen-tts
qwen-tts-demo --help
qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-Base --ip 127.0.0.1 --port 8000Open http://127.0.0.1:8000 on the same computer. The initial run downloads model files. Choose a saved reference recording, enter its exact transcript, and then the new sentence you want spoken. For a built-in speaker instead, load a CustomVoice checkpoint.
The CUDA examples are not Mac installation commands. FlashAttention is an optional acceleration dependency with its own GPU and precision requirements; it is not something to install blindly on Apple Silicon. Keep the backend's supported dependency versions together. Once all assets are cached, test a second run with the network disconnected before relying on offline operation.
Mac and Windows desktop workflow
Local AI Audio uses the speech module of TopLocal Studio. Its model catalog selects a 0.6B Base Q8 GGUF through audio.cpp, rather than the official 1.7B Python example. The catalog lists roughly 2 GB of model downloads; that is disk usage, not peak working memory.
- Open the speech module and check the available Qwen3-TTS model in your installed version.
- Download the model and review its license.
- Add one clear speaker recording and the text it contains.
- Generate a short sentence first, listen for skipped words and pronunciation, then extend the script.
- Save the result locally. Repeat after disconnecting the network if offline operation matters to your workflow.
Mac builds use Apple Silicon; the Windows desktop route uses the packaged backend, not the CUDA recipe above. We have not benchmarked the current Windows installer in this article. Check the app's platform and download details before choosing an installer.
What we measured on a Mac
Download the historical measurement summary CSV. This is an extract of the September 23 record, not a new benchmark.
Test setup: MacBook Pro M5 Pro, 64 GB unified memory, audio.cpp commit 487800f5, 0.6B Q8 GGUF, September 23, 2026. Metal and CPU were tested separately.
| Measurement | Recorded result |
|---|---|
| Metal generation speed | 2.6–3.4 times audio playback speed |
| CPU generation speed | About 1 times playback speed |
| Reported working memory | 3.1–3.6 GB |
| English / mixed Chinese-English ASR round-trip error | 0.0% on the selected short passages |
| Chinese / Cantonese round-trip error | 1.9% / 4.2% on the selected passages |
These are small-sample measurements on one computer. The error score came from transcribing generated speech with Qwen3-ASR-1.7B and comparing text. It does not measure speaker similarity or guarantee error-free speech. The reference clips were project example recordings, not a new voice recording for this article. These memory results do not establish a minimum for every Mac, NVIDIA card, or longer script.
One mixed-language input was: “今天的 meeting 改到下午三点,请把最新的 PPT 发到我的 email,谢谢。” Listen for whether the English words survive, not just whether the voice sounds pleasant.
Hear the existing voice-cloning example
These are the existing product demonstration clips, separate from the benchmark above. They illustrate the reference → new text workflow; they are not a speed comparison.
Reference: 很久很久以前,在一片安静的森林里,住着一只爱看星星的小狐狸。每天晚上,它都会爬上山坡,数着天上的星星慢慢睡去。
New text: 这是用刚才那段录音克隆出来的声音,我可以用它朗读任何新的内容,比如这句话。
Fix the common mismatches
- No cloning controls: confirm you loaded Base, not CustomVoice or VoiceDesign.
- Words missing or pronunciation wrong: shorten the sentence, write numbers as intended, and check the reference transcript. A better voice match cannot fix incorrect input text.
- Out of memory: lower batch size and generation length, close other model sessions, then consider a smaller checkpoint or backend-supported quantization. Do not rename an official weight file to
.gguf. - Cannot record in the browser: upload a saved file; microphone permissions depend on browser and HTTPS settings.
- Still contacting the network: check whether model, tokenizer or other runtime files remain uncached. A locally displayed web UI alone does not prove offline inference.
Use your own voice or a reference you have permission to use. For recording advice, read offline voice cloning. For the opposite task—audio into text and subtitles—use offline transcription, not a TTS checkpoint. Comparing hosted and desktop products? See the offline ElevenLabs alternative.
More Posts

Offline Text-to-Speech, Voice Cloning and Transcription with Subtitles
How to turn recordings into text and SRT subtitles, and text into natural speech, entirely on your own Mac or Windows PC. Covers model choice, speed, voice cloning and responsible use.

MacWhisper and Superwhisper Alternatives for Transcribing Audio and Video Files
MacWhisper transcribes files, Superwhisper is mainly for dictation. What each does, how they differ, and an offline alternative for Mac and Windows that turns recordings into text and SRT subtitles.

How to Transcribe Audio to Text Offline (Audio and Video to Text Converter with Subtitles)
Transcribe audio and video to text on your own Mac or Windows PC. Get a TXT transcript and SRT subtitles with real timings, with no uploads and no length limit.
Newsletter
Join the community
Subscribe to our newsletter for the latest news and updates