Run Qwen3-TTS Locally: Setup, Voice Cloning and Hardware
2026/10/05

Run Qwen3-TTS Locally: Setup, Voice Cloning and Hardware

Run Qwen3-TTS locally with official Python tools or a desktop workflow. Compare checkpoints, Mac benchmark results, reference audio and setup pitfalls.

Qwen3-TTS has downloadable models for local speech generation. Choose Base for reference-audio cloning, CustomVoice for built-in speakers, or VoiceDesign for describing a new voice. Those are different checkpoints, not interchangeable modes of the same file. This guide covers the official setup and the smaller quantized model used by Local AI Audio.

Checked on October 5, 2026. Performance numbers below come from our September 23 test, not a new test of every current installer. Qwen-Audio-3.1 is a different model family: see cloud APIs versus local Qwen3-TTS.

Pick a route and checkpoint

GoalModel / routeWhat to check first
Read in a built-in voiceOfficial CustomVoice checkpointSupported speaker names and languages
Clone a reference voiceOfficial 0.6B or 1.7B Base checkpointClean reference recording and matching transcript
Describe a new speakerOfficial 1.7B VoiceDesign checkpointA separate model download
Use the desktop speech workflowLocal AI Audio's Qwen3-TTS 0.6B Base Q8 via audio.cppModel availability in your installed release

The official repository documents the Python route. The 0.6B Base model card identifies the smaller cloning checkpoint and its Apache-2.0 license. A downloaded weight file is not the whole runtime: keep the tokenizer and dependencies required by your chosen backend.

Official local setup

Use a fresh environment with a CUDA-compatible PyTorch installation for the documented NVIDIA examples. The upstream quickstart recommends Python 3.12:

conda create -n qwen3-tts python=3.12 -y
conda activate qwen3-tts
python -m pip install -U qwen-tts
qwen-tts-demo --help
qwen-tts-demo Qwen/Qwen3-TTS-12Hz-1.7B-Base --ip 127.0.0.1 --port 8000

Open http://127.0.0.1:8000 on the same computer. The initial run downloads model files. Choose a saved reference recording, enter its exact transcript, and then the new sentence you want spoken. For a built-in speaker instead, load a CustomVoice checkpoint.

The CUDA examples are not Mac installation commands. FlashAttention is an optional acceleration dependency with its own GPU and precision requirements; it is not something to install blindly on Apple Silicon. Keep the backend's supported dependency versions together. Once all assets are cached, test a second run with the network disconnected before relying on offline operation.

Mac and Windows desktop workflow

Local AI Audio uses the speech module of TopLocal Studio. Its model catalog selects a 0.6B Base Q8 GGUF through audio.cpp, rather than the official 1.7B Python example. The catalog lists roughly 2 GB of model downloads; that is disk usage, not peak working memory.

  1. Open the speech module and check the available Qwen3-TTS model in your installed version.
  2. Download the model and review its license.
  3. Add one clear speaker recording and the text it contains.
  4. Generate a short sentence first, listen for skipped words and pronunciation, then extend the script.
  5. Save the result locally. Repeat after disconnecting the network if offline operation matters to your workflow.

Mac builds use Apple Silicon; the Windows desktop route uses the packaged backend, not the CUDA recipe above. We have not benchmarked the current Windows installer in this article. Check the app's platform and download details before choosing an installer.

What we measured on a Mac

Download the historical measurement summary CSV. This is an extract of the September 23 record, not a new benchmark.

Test setup: MacBook Pro M5 Pro, 64 GB unified memory, audio.cpp commit 487800f5, 0.6B Q8 GGUF, September 23, 2026. Metal and CPU were tested separately.

MeasurementRecorded result
Metal generation speed2.6–3.4 times audio playback speed
CPU generation speedAbout 1 times playback speed
Reported working memory3.1–3.6 GB
English / mixed Chinese-English ASR round-trip error0.0% on the selected short passages
Chinese / Cantonese round-trip error1.9% / 4.2% on the selected passages

These are small-sample measurements on one computer. The error score came from transcribing generated speech with Qwen3-ASR-1.7B and comparing text. It does not measure speaker similarity or guarantee error-free speech. The reference clips were project example recordings, not a new voice recording for this article. These memory results do not establish a minimum for every Mac, NVIDIA card, or longer script.

One mixed-language input was: “今天的 meeting 改到下午三点,请把最新的 PPT 发到我的 email,谢谢。” Listen for whether the English words survive, not just whether the voice sounds pleasant.

Hear the existing voice-cloning example

These are the existing product demonstration clips, separate from the benchmark above. They illustrate the reference → new text workflow; they are not a speed comparison.

Reference: 很久很久以前,在一片安静的森林里,住着一只爱看星星的小狐狸。每天晚上,它都会爬上山坡,数着天上的星星慢慢睡去。

New text: 这是用刚才那段录音克隆出来的声音,我可以用它朗读任何新的内容,比如这句话。

Fix the common mismatches

  • No cloning controls: confirm you loaded Base, not CustomVoice or VoiceDesign.
  • Words missing or pronunciation wrong: shorten the sentence, write numbers as intended, and check the reference transcript. A better voice match cannot fix incorrect input text.
  • Out of memory: lower batch size and generation length, close other model sessions, then consider a smaller checkpoint or backend-supported quantization. Do not rename an official weight file to .gguf.
  • Cannot record in the browser: upload a saved file; microphone permissions depend on browser and HTTPS settings.
  • Still contacting the network: check whether model, tokenizer or other runtime files remain uncached. A locally displayed web UI alone does not prove offline inference.

Use your own voice or a reference you have permission to use. For recording advice, read offline voice cloning. For the opposite task—audio into text and subtitles—use offline transcription, not a TTS checkpoint. Comparing hosted and desktop products? See the offline ElevenLabs alternative.

Newsletter

Join the community

Subscribe to our newsletter for the latest news and updates