Swedish Text to Speech: Free, In-Browser, Nothing Uploaded
Swedish sits in the same one-engine spot as Korean and Japanese, with the same tradeoff: the good voice needs a GPU. Paste Swedish text, press play, and the audio is generated on your own device — no account, no upload, no character limit. This page is the honest version: which engines really speak Swedish (Svenska), which ones silently do not, and what each one costs you.
Which engines actually speak Swedish
Quick TTS ships four TTS engines and lets you switch between them from one dropdown — but not all four cover every language, and the gaps are not the ones you would guess from the marketing.
| Engine | Swedish? | What that means |
|---|---|---|
| Web Speech API | Yes | Your operating system's built-in Swedish voices. No download, works on every device including iPhone and Android. |
| Piper (WebAssembly) | No | Quick TTS does not ship a Piper Swedish voice yet — upstream carries one, but it has not passed our intelligibility checks — so the option is hidden rather than offered and then failing. |
| Kokoro-82M (WebGPU) | No | English only. Hidden on every other language, because a voice you cannot select is not an option. |
| Supertonic HD (WebGPU) | Yes | 44.1kHz studio-grade Swedish, one of thirty-one mapped languages. WebGPU-only, ~380MB cached once. |
Only one of Quick TTS's three neural engines speaks Swedish, and it is worth being blunt about the other two:
- Piper ships zero Swedish voices. Quick TTS does not ship a Piper Swedish voice in this build — the upstream catalogue does carry one, but it has not been through our per-voice intelligibility checks yet — so the engine option is hidden here rather than shown and then failing to load. The Piper option is hidden on Swedish rather than offered and then failing.
- Kokoro-82M is English-only and is hidden everywhere else.
- Supertonic HD does speak Swedish, natively, at 44.1kHz, as one of the thirty-one languages this build maps.
That leaves exactly one neural option in the browser for Swedish, and it happens to be the best one of the three — the same position Korean and Japanese are in. The Web Speech API fills the gap everywhere Supertonic can't reach.
The honest limits
Being precise about that gap: Supertonic is WebGPU-only, with no WebAssembly fallback tier, and there is no Piper voice sitting underneath it the way there is for Spanish or Russian. On a desktop with a working GPU you get studio-grade 44.1kHz Swedish after a one-time ~380MB cached download. On a phone, in Safari, or on a machine with GPU acceleration disabled, you get your operating system's Swedish voice and nothing in between — there is no neural middle tier for Swedish the way Piper provides one for languages it does cover.
That floor is not bad by browser-TTS standards: Windows ships a classic voice, Edge layers Azure-quality Online voices (Sofie, Mattias) on top, and macOS/iOS ship Alva and Oskar. So the honest recommendation for most visitors on this page is not "wait for WebGPU" — it's "use Edge's Online voices on desktop, or your phone's built-in Swedish voice, and reach for Supertonic HD only if you already have a WebGPU-capable machine."
Swedish also gives every engine the same specific problem: pitch accent. Word pairs like anden (the duck) and anden (the spirit) are spelled identically but distinguished only by tone, and an engine trained mostly on tone-flat languages will sometimes collapse that distinction. Long closed compounds (sjukhusdirektör, arbetsmarknadsstyrelsen) are the other recurring stumble — Swedish writes multi-word concepts as one unbroken word, and getting the internal stress pattern right across a five- or six-morpheme compound is harder than reading the same concept written as separate words would be.
Nothing you paste is uploaded
This is the part that separates in-browser TTS from the cloud tools that look identical in a browser tab. A cloud tool sends your Swedish text to a server, synthesizes it there, and sends audio back — the page is a remote control. Quick TTS ships the engine to you instead and runs it on your CPU or GPU, so your text is never received by anyone. The test is simple: once the page and its model have loaded, the neural engines keep working with your network disconnected.
That matters most for the documents people actually want read aloud — a contract, a medical letter, coursework, an unpublished draft. It also means there is no "we phonemize non-English text on our server" asterisk, which is not true of every tool advertising local synthesis.
Reading Swedish documents, not just pasted text
PDF, DOCX, EPUB, ODT, RTF, HTML, TXT and Markdown are all parsed in the browser, so a whole Swedish EPUB or a long PDF opens and plays without a byte leaving your machine. A scanned PDF — a photo of text with no text layer — is handled by running OCR locally in WebAssembly rather than uploading the scan, which is precisely the case where local processing is worth the most.
Speed, volume and voice can all be changed mid-read and the current passage replays under the new settings without starting over. Switching the engine itself stops playback, so you press play once to start the new one — there is no mid-read handover of the remaining text between engines.
Frequently asked questions
Is Swedish text-to-speech free?
Yes. Quick TTS reads Swedish with no account, no character limit and no watermark. Synthesis runs in your browser on your own device, so there is no per-character server cost to meter and nothing to gate behind a sign-up.
Is my Swedish text uploaded anywhere?
No. The speech engine ships to your browser and runs locally, so the text you paste is turned into audio on your machine and never sent to a server. That is an architectural property, not a promise in a privacy policy — there is no server in the loop once the page has loaded.
What is the best Swedish TTS voice in the browser?
Supertonic HD is the highest-quality Swedish voice available here — 44.1kHz, running on WebGPU. It is also the only neural Swedish option: Piper ships no Swedish voice and Kokoro is English-only, so the alternative is your operating system's built-in voice. To get Swedish voices rather than English ones, open the app at quick-tts.com/sv/ — the neural voice list follows the page language.
Does it work on a phone?
The Web Speech API engine works on every modern phone with no download, and it is the default so the first play click always works. The neural engines are more demanding: Supertonic HD needs WebGPU and is effectively a desktop feature.
Can I download the Swedish audio?
Yes, as WAV or MP3, generated on your device. Anything longer than about 20,000 characters is packaged as a ZIP of audio parts so memory stays bounded on long documents.
Try it in Swedish — open the Svenska app, not the English one
This is the one piece of setup worth knowing, because it is not obvious: the neural voice list follows the language of the page you are on. On the English homepage the Supertonic reads with the English language tag — so Swedish text pasted there will be pronounced as though it were English. Open the Svenska version instead and the engine picks up Swedish automatically.
The one exception is Supertonic's "Auto language" toggle, which switches generation to a language-agnostic mode that reads mixed-language text with the right pronunciation per language without tagging anything by hand. That works from any page — it is the right choice for a document that is part Swedish and part English.
For the architecture behind all of this, see how in-browser text-to-speech works; for a tool-by-tool comparison with the other local TTS sites, see the comparison page.