← All languages

Estonian Text to Speech: Free, In-Browser, Nothing Uploaded

Estonian sits in the same one-engine spot as Korean and Hindi: exactly one neural voice speaks it, and it needs a GPU most Estonian readers on Windows will not have a voice to fall back to without. Paste Estonian text, press play, and the audio is generated on your own device — no account, no upload, no character limit. This page is the honest version: which engines really speak Estonian (Eesti), which ones silently do not, and what each one costs you.

Which engines actually speak Estonian

Quick TTS ships four TTS engines and lets you switch between them from one dropdown — but not all four cover every language, and the gaps are not the ones you would guess from the marketing.

EngineEstonian?What that means
Web Speech APIYesYour operating system's built-in Estonian voices. No download, works on every device including iPhone and Android.
Piper (WebAssembly)NoQuick TTS does not ship a Piper Estonian voice yet — upstream carries one, but it has not passed our intelligibility checks — so the option is hidden rather than offered and then failing.
Kokoro-82M (WebGPU)NoEnglish only. Hidden on every other language, because a voice you cannot select is not an option.
Supertonic HD (WebGPU)Yes44.1kHz studio-grade Estonian, one of thirty-one mapped languages. WebGPU-only, ~380MB cached once.

Only one of Quick TTS's three neural engines speaks Estonian, and it is worth being blunt about the other two:

That is the same one-engine position Korean and Hindi are in. The difference for Estonian is what sits underneath: a default Windows install frequently ships no classic Estonian voice at all, so on that platform the Web Speech API tier that other languages fall back to safely does not reliably exist. Edge's online neural voices and Android's Google TTS are the practical alternatives when the GPU tier is unavailable.

The honest limits

Being precise about that, because it is the whole picture for Estonian: Supertonic is WebGPU-only with no WebAssembly fallback tier, and there is no Piper voice sitting underneath it the way there is for Spanish or Russian. On a desktop with a working GPU you get studio-grade 44.1kHz Estonian after a one-time ~380MB cached download. On a phone, or on a Windows machine without a working GPU, you get whatever your operating system ships, and nothing in between.

The Web Speech tier is genuinely weaker here than for most languages this site covers. Apple does not ship an official Estonian system voice on macOS or iOS. A default Windows installation commonly has no classic Estonian voice either, though Microsoft Edge exposes two online neural voices (Anu and Kert) that fill that gap when a network connection is available. Android has carried Estonian in Google's text-to-speech engine since 2018, which makes it the most reliable non-GPU floor for this language. There is no way to dress this up: Estonian's non-neural coverage is thinner than average, and that is the reason this page exists rather than a page claiming Estonian is fully covered everywhere.

Estonian's three-way distinction between short, long, and "overlong" vowel and consonant length is only partly visible in ordinary spelling, since all three degrees can share the same written letters. A synthetic voice trained mostly on the more common short and long contrast can flatten the overlong distinction that changes a word's meaning, so an odd-sounding vowel or consonant length in a generated sentence is worth checking against the source text rather than assumed to be a mispronunciation.

Nothing you paste is uploaded

This is the part that separates in-browser TTS from the cloud tools that look identical in a browser tab. A cloud tool sends your Estonian text to a server, synthesizes it there, and sends audio back — the page is a remote control. Quick TTS ships the engine to you instead and runs it on your CPU or GPU, so your text is never received by anyone. The test is simple: once the page and its model have loaded, the neural engines keep working with your network disconnected.

That matters most for the documents people actually want read aloud — a contract, a medical letter, coursework, an unpublished draft. It also means there is no "we phonemize non-English text on our server" asterisk, which is not true of every tool advertising local synthesis.

Reading Estonian documents, not just pasted text

PDF, DOCX, EPUB, ODT, RTF, HTML, TXT and Markdown are all parsed in the browser, so a whole Estonian EPUB or a long PDF opens and plays without a byte leaving your machine. A scanned PDF — a photo of text with no text layer — is handled by running OCR locally in WebAssembly rather than uploading the scan, which is precisely the case where local processing is worth the most.

Speed, volume and voice can all be changed mid-read and the current passage replays under the new settings without starting over. Switching the engine itself stops playback, so you press play once to start the new one — there is no mid-read handover of the remaining text between engines.

Frequently asked questions

Is Estonian text-to-speech free?

Yes. Quick TTS reads Estonian with no account, no character limit and no watermark. Synthesis runs in your browser on your own device, so there is no per-character server cost to meter and nothing to gate behind a sign-up.

Is my Estonian text uploaded anywhere?

No. The speech engine ships to your browser and runs locally, so the text you paste is turned into audio on your machine and never sent to a server. That is an architectural property, not a promise in a privacy policy — there is no server in the loop once the page has loaded.

What is the best Estonian TTS voice in the browser?

Supertonic HD is the highest-quality Estonian voice available here — 44.1kHz, running on WebGPU. It is also the only neural Estonian option: Piper ships no Estonian voice and Kokoro is English-only, so the alternative is your operating system's built-in voice. To get Estonian voices rather than English ones, open the app at quick-tts.com/et/ — the neural voice list follows the page language.

Does it work on a phone?

The Web Speech API engine works on every modern phone with no download, and it is the default so the first play click always works. The neural engines are more demanding: Supertonic HD needs WebGPU and is effectively a desktop feature.

Can I download the Estonian audio?

Yes, as WAV or MP3, generated on your device. Anything longer than about 20,000 characters is packaged as a ZIP of audio parts so memory stays bounded on long documents.

Try it in Estonian — open the Eesti app, not the English one

This is the one piece of setup worth knowing, because it is not obvious: the neural voice list follows the language of the page you are on. On the English homepage the Supertonic reads with the English language tag — so Estonian text pasted there will be pronounced as though it were English. Open the Eesti version instead and the engine picks up Estonian automatically.

The one exception is Supertonic's "Auto language" toggle, which switches generation to a language-agnostic mode that reads mixed-language text with the right pronunciation per language without tagging anything by hand. That works from any page — it is the right choice for a document that is part Estonian and part English.

For the architecture behind all of this, see how in-browser text-to-speech works; for a tool-by-tool comparison with the other local TTS sites, see the comparison page.