Offline text to speech & speech to text · Mac and Windows
Transcribe audio to text. Give text a voice.
Turn hours of recordings and video into text and SRT subtitles, make natural voiceovers with offline text to speech, and clone a voice from a few seconds of audio. All on your own computer.
- 100% offline
- No subscription or credits
- Mac and Windows
- Open source
Upload, transcribe, download subtitles
Voices
Eight voices, ready to read.
Preset Chinese voices for narration, news, ads and stories. Press play.
Voice cloning
Voice cloning from a few seconds of audio.
Record 5 to 20 seconds, and the app reads any new text in that voice.
Audio to text converter
Transcripts and subtitles for any length.
Drop in a recording or a video. Get the text and an SRT subtitle file, with dialects and accents handled.
More samples
Speech to text
A 15-second recording, transcribed on this computer:
1 00:00:00,136 --> 00:00:02,496 大家好,欢迎使用本地创作台。 2 00:00:02,496 --> 00:00:08,226 今天我想和大家聊聊怎样在自己的电脑上用人工智能写歌、画画和剪辑视频。 3 00:00:08,226 --> 00:00:11,004 所有内容都在本机生成,不需要联网, 4 00:00:11,004 --> 00:00:12,312 也不用担心隐私。
Models
The right model for each job.
Choose per task inside the app. All run offline.
Qwen3-ASR 0.6B
RecommendedAccurate and fast transcription in Mandarin, English and more.
about 3 min per hour of audio · 1.2 GB
Qwen3-ASR 1.7B
Better with dialects and strong accents.
about 5 min per hour · 2.5 GB
SenseVoice
The fastest, for very long recordings.
about 1 min per hour · 0.25 GB
Kokoro
RecommendedFast voiceovers with preset Chinese voices.
a few seconds · 0.2 GB
Qwen3-TTS
More natural reading, mixed Chinese and English, voice cloning.
10–20 seconds · 2 GB
Why do speech locally?
Any length
Hours of audio are fine; nothing is uploaded or metered.
Subtitles included
Every transcript comes with an SRT file with real timings.
Voice cloning
A few seconds of your voice is enough to read new text.
Chinese and English
Mixed-language text reads naturally.
Private
Recordings of meetings and interviews stay on your computer.
Video in, text out
Drop an mp4 or mov and get its transcript.
Three steps.
- 01
Download
Get the installer for macOS or Windows.
- 02
Pick a model
The app suggests models that fit your computer and downloads them for you.
- 03
Create offline
Type an idea or click an example. Results stay on your disk.
Frequently asked questions
Can it generate subtitles from a video?+
Yes. It works as a subtitle generator: drop in an mp4 or mov file and you get a TXT transcript plus an SRT subtitle file with timings, ready for your video editor.
Is this a Whisper alternative?+
It does the same job with different models. The app does not use Whisper; it uses Qwen3-ASR (0.6B and 1.7B) and SenseVoice Small, which are especially strong on Mandarin and Cantonese and also handle English, Japanese and Korean.
Which languages are supported?+
Transcription handles Mandarin, Cantonese, English, Japanese, Korean and more, and can convert Traditional Chinese output to Simplified. For text to speech, Kokoro has preset Chinese and English voices, and Qwen3-TTS reads mixed Chinese and English.
How accurate is the transcription?+
Qwen3-ASR is among the most accurate open models for Mandarin; the 1.7B model is more robust with dialects and accents.
Is voice cloning allowed for anyone's voice?+
Only clone voices you have permission to use. The app processes everything locally and keeps nothing outside your computer.
Does it really work offline?+
Yes. After the models are downloaded once, everything runs on your computer without an internet connection. Nothing is uploaded.
Are the AI models included in the installer?+
No. The installer contains the app and its runtime. You download only the models you want from inside the app, and each one shows its license first.
Is it open source?+
Yes, the app is open source under Apache-2.0 and you can build it yourself. The ready-made installer is available through the download button.
Model guides and practical workflows
Check the model, hardware and license before you install.
Qwen3-TTS local setup
Choose the right checkpoint, prepare a reference voice and compare measured Mac results.
Qwen Audio 3.1 vs Qwen3-TTS
Separate hosted audio APIs from models you can download and run locally.
Offline transcription and subtitles
Turn recordings into text and SRT files with a local speech recognition model.