← Back to Quick TTS

Text to Speech by Language

Most browser TTS tools ship English and stop. Quick TTS ships four engines across sixteen locales — but the coverage is uneven, and which engine speaks your language is not obvious from any dropdown. These pages say plainly which engines work, which are missing, and why.

Coverage at a glance

LanguagePiper (no GPU needed)Supertonic HDThe short version
Korean 한국어 Yes — 44.1kHz HD only — no Piper voice exists
Japanese 日本語 Yes — 44.1kHz HD only — no Piper voice exists
Chinese 中文 zh_CN-huayan-medium Piper only — no Supertonic Chinese
Spanish Español es_ES-davefx-medium +1 Yes — 44.1kHz Both neural tiers
Russian Русский ru_RU-irina-medium Yes — 44.1kHz Both neural tiers
Vietnamese Tiếng Việt vi_VN-vais1000-medium Yes — 44.1kHz Both neural tiers

Kokoro-82M is missing from that table on purpose: it synthesizes English only and is hidden on every other language, so listing it per-language would only ever mean "no".

Why the gaps run in both directions

It would be tidier if one engine were simply best everywhere. It is not. Supertonic HD is the highest-quality engine and it is the only neural option for Korean and Japanese — but it ships no Chinese at all, so Mandarin runs on Piper, the engine that is elsewhere the compromise choice. Piper covers thirteen locales and needs no GPU; Supertonic covers fifteen and needs WebGPU with no WebAssembly fallback. Neither is a superset of the other.

The practical consequence: the engine worth using depends on your language and your hardware, in that order. Each page below states both.

Pick a language

Every one of them runs in the browser with nothing uploaded, no account, and no character limit. If you want the architecture rather than the language detail, how in-browser text-to-speech works covers it, and the comparison page puts Quick TTS next to the other local-first TTS tools.