← Quick TTS

Every Piper voice Quick TTS ships

Piper is the most-selected engine on this site, and it is the one that runs without a GPU. These are the 16 voices you can actually choose, across 13 languages, each with its own page covering the quality tier, what happens when your device cannot load it, and what the alternative on that language is. Everything runs in your browser; nothing you paste is uploaded.

The voices

VoiceLanguageTierRole Step-down?HD alternative?
ar_JO-kareem-medium Arabic medium Default Yes Yes
de_DE-thorsten-medium German medium Default No Yes
en_GB-vctk-medium English medium Alternative No Yes
en_US-joe-medium English medium Alternative No Yes
en_US-libritts_r-medium English medium Default Yes Yes
es_ES-davefx-medium Spanish medium Default Yes Yes
es_MX-claude-high Spanish high Alternative No Yes
fr_FR-siwis-medium French medium Default No Yes
it_IT-riccardo-x_low Italian x_low Default No Yes
nl_NL-mls-medium Dutch medium Default Yes Yes
pl_PL-gosia-medium Polish medium Default Yes Yes
pt_BR-faber-medium Brazilian Portuguese medium Default No Yes
ru_RU-irina-medium Russian medium Default No Yes
tr_TR-fettah-medium Turkish medium Default No Yes
vi_VN-vais1000-medium Vietnamese medium Default No Yes
zh_CN-huayan-medium Mandarin Chinese medium Default No No

By language

The voices you will never see in the dropdown

Quick TTS also carries 5 smaller Piper models that are never selectable: they do not appear in the voice list and cannot be chosen by name. They exist only as the automatic landing place when a device cannot allocate a session for the larger model above them:

They are listed here because the alternative is worse: a reader who sees one of these IDs in a log or an error message should be able to find out what it is, without us implying it is a voice they can pick.

Which one should you use?

On every language except English and Spanish there is exactly one choice, so the real question is which engine to use rather than which voice. Piper is the right answer when you have no usable GPU, and it is the only neural option at all for Mandarin Chinese. Where Supertonic HD covers the language, it is the higher-quality option and it needs WebGPU, which in practice means a desktop.

Speed, volume and voice can all be changed mid-read and the current passage replays under the new settings without starting over. Switching the engine itself stops playback, so you press play once to start the new one. There is no mid-read handover of the remaining text between engines.

Nothing you paste is uploaded

Piper runs as WebAssembly inside your own browser tab. The model is downloaded once and cached; after that, the text you paste is turned into audio on your machine and never sent anywhere. That is an architectural property rather than a policy promise: the test is that generation keeps working with your network disconnected once the model has loaded.

Read next

Try them now →