Browser text-to-speech: hear every voice your browser has built in
Browser text-to-speech, or browser TTS, is speech a web page makes with the voices already built into your browser and device, through the Web Speech API’s speechSynthesis. The tester below lists every voice your browser reports, with its language, whether it runs on your device or on a remote service, and which one is the default, and plays a voice sample from any of them.
This page sends what you type nowhere. A voice marked remote does: when it speaks, the browser sends the text to that voice’s speech service. To speak in your own voice instead, which no browser can do by itself, skip to using a cloned voice in a browser.
VOICES YOUR BROWSER REPORTS
Reading the voice list from your browser…
What browser TTS is
The Web Speech API has two halves: speech recognition, which turns speech into text, and speech synthesis, which turns text into speech. Browser TTS is the second. A page describes what to say with a SpeechSynthesisUtterance (the text, the voice, the language, the speed and the pitch) and hands it to speechSynthesis.speak(), which queues it and reads it out through the device’s speakers. speechSynthesis.getVoices() lists the voices available on the device, and speechSynthesis.cancel() empties the queue and stops speaking at once.
Why the voice list differs by browser and device
A page ships no voices of its own. It gets the ones the browser offers, and every browser offers a different set: most systems have a speech synthesizer of their own, browsers pass its voices through, and some browsers add voices that run on a remote service. So the same page lists different voices in Edge on Windows, Chrome on a Mac and Firefox on Linux. Each voice says which kind it is.
- Local voices (
localServiceis true) are spoken by the synthesizer on your device. The text stays on the device. - Remote voices (
localServiceis false) are spoken by a service over the network. The text you play with one is sent to that service, and MDN notes that remote voices can add latency, bandwidth or cost. - Microsoft Edge’s “Online (Natural)” voices, with names like “Microsoft Guy Online (Natural) - English (United States)” and Ana, Eric and many more beside it, are Microsoft’s neural voices, and they are remote: the text is sent to Microsoft’s speech service. Windows also names its own installed voices after Microsoft, and those run on the device.
- Chrome’s “Google” voices, such as “Google US English”, are voices Chrome adds on the desktop next to the system’s own, and Chrome reports them as remote.
None of this has to be taken on trust: the tester above shows the name, language, local or remote flag and default that your own browser reports, voice by voice.
Browser TTS in a few lines of JavaScript
No library, key or account: this runs as it is in the console of any browser that has the Web Speech API. Listen for voiceschanged, because in Chrome the list is empty until the voices have loaded.
// List the voices this browser reports. The list can start empty and fill in later.
function listVoices() {
for (const voice of speechSynthesis.getVoices()) {
const where = voice.localService ? "local" : "remote";
console.log(voice.name, voice.lang, where, voice.default ? "default" : "");
}
}
listVoices();
speechSynthesis.addEventListener("voiceschanged", listVoices);
// Speak a sentence, in the first English voice if there is one.
const utterance = new SpeechSynthesisUtterance("Hello from your browser.");
const english = speechSynthesis
.getVoices()
.find((voice) => voice.lang.startsWith("en"));
utterance.voice = english ?? null; // null: the browser's default voice
utterance.rate = 1; // 0.1 to 10
utterance.pitch = 1; // 0 to 2
speechSynthesis.speak(utterance);
// Stop speaking and empty the queue.
speechSynthesis.cancel();The limits of browser TTS
- You cannot choose your visitors’ voices. A page picks from the list each visitor’s browser reports, and that list can be long, short or empty.
- There is no audio file. Speech goes straight to the device’s audio output. The API hands the page no audio to save, upload or download, and none of an utterance’s events (start, end, boundary, mark, pause, resume, error) carries any.
- A page cannot add a voice. Nothing a web page does installs a voice into
speechSynthesis, a cloned voice included. - Speed and pitch are requests. The API takes a rate from 0.1 to 10 and a pitch from 0 to 2, and a voice may hold either to a narrower range.
- Remote voices send your text away to the service that speaks it, and need the network to speak at all.
- The list can load late. In Chrome it is empty until
voiceschangedfires, so code that reads it once at page load can find nothing.
Use your own cloned voice in a browser
Because a page cannot add a voice to the browser, your own voice has to arrive as audio. Clone it with a text-to-speech service, have the service speak the text, and play the result in the page with an <audio> element or new Audio(url). That also gets you the audio file browser TTS never gives you.
With VoiceLabs you clone a voice from a 2–30 second recording of clear speech, either in the VoiceLabs studio in your browser or through the API (POST /v1/voices/clone, or cloneVoice() in the @voicelabs/sdk package). Clone only a voice you have the right to use: the responsible use policy says what that means. After that, a request with your text and the voice’s name comes back as a take in that voice.
Keep the API key on a server. An API key spends your account’s audio, and the SDK’s own documentation puts the rule plainly:
“In a browser, only ever use a key you are willing to make public.”
So the page asks your own server for the speech, and the server, which holds the key, asks VoiceLabs.
A server route that speaks in your cloned voice
A Next.js route handler with the SDK. generateSpeech() starts the generation and waits for it to finish; the audio_url it comes back with is a signed, time-limited link, so the browser can play it without ever seeing the key.
// app/api/speak/route.ts: runs on your server, so the API key never reaches a browser.
import { VoiceLabs } from "@voicelabs/sdk";
const voicelabs = new VoiceLabs({ apiKey: process.env.VOICELABS_API_KEY! });
export async function POST(request: Request) {
const { text } = (await request.json()) as { text: string };
// In a real app, check who is asking and cap the text first.
const generation = await voicelabs.generateSpeech({
text,
voice_name: "My voice", // the name you gave your cloned voice
});
// A signed, time-limited link to the audio: a browser can play it without the key.
return Response.json({ audioUrl: generation.audio_url });
}And the page that plays it. Browsers can refuse to start sound the visitor did not ask for, so run it from a click, and keep the element’s controls so they can press play themselves.
// In the page, from a click handler.
const response = await fetch("/api/speak", {
method: "POST",
headers: { "content-type": "application/json" },
body: JSON.stringify({ text: "Hello, in my own voice." }),
});
if (!response.ok) throw new Error(`Speech failed with HTTP ${response.status}`);
const { audioUrl } = await response.json();
const player = document.querySelector("audio"); // an <audio controls> element
player.src = audioUrl;
await player.play();That route answers anyone who can reach it. A real one checks who is asking and caps the text, because every request spends your plan’s audio. The API is part of VoiceLabs Pro: on an account without an active subscription or trial, every request to it is refused. Authentication, errors and rate limits are in the API guide.
Questions about browser text-to-speech
- What is browser TTS?
- Browser TTS (browser text-to-speech) is speech a web page makes with the voices the browser already has, through the speech synthesis half of the Web Speech API. The page puts the text in a SpeechSynthesisUtterance, picks one of the voices speechSynthesis.getVoices() lists, and calls speechSynthesis.speak(), and the browser reads it out through the device's speakers. Nothing has to be installed, but which voices a visitor hears depends on their browser and device.
- Why are my browser's voices called Microsoft Ana, Eric or Guy?
- Because they are Microsoft's voices, and you are most likely using Microsoft Edge. Edge lists Microsoft's online neural voices, with names like "Microsoft Guy Online (Natural) - English (United States)", alongside the voices installed on the device. The Online (Natural) voices are remote: the browser reports them with localService set to false, and the text you play with one is sent to Microsoft's speech service to be spoken. Windows also names its own installed voices after Microsoft, and those run on the device. The tester on this page shows which of your voices are local and which are remote.
- Why is my voice list empty?
- Usually because the browser is still loading it. Chrome fills the list after the page loads and announces it with the voiceschanged event, so code that reads speechSynthesis.getVoices() once, too early, gets nothing back; the tester on this page listens for that event. If the list stays empty, the browser has no voices to offer on that device: some systems, Linux desktops without a speech synthesizer among them, report none. A browser without the Web Speech API cannot speak text by itself at all.
- Can I clone my voice so the browser can use it?
- Not as a browser voice. A web page can list, choose and speak with the voices the browser reports, but it cannot add one, so a cloned voice can never appear in speechSynthesis.getVoices(). What works instead is to make the speech with a voice-cloning text-to-speech service and play it in the page as ordinary audio, with an audio element or new Audio(url). With VoiceLabs you clone a voice from a 2–30 second recording, in the studio or through the API, and send text to be spoken in it. Clone only a voice you have the right to use.The API guideResponsible use
- Can I download browser TTS as an audio file?
- No. The Web Speech API plays speech straight to the device's audio output and gives the page no audio back: speechSynthesis.speak() returns nothing, and the events an utterance fires (start, end, boundary, mark, pause, resume and error) carry no sound. So a page cannot save what a browser voice says as a file. To get a file, the speech has to come from a text-to-speech service that returns audio, the way the VoiceLabs API returns each finished take as an audio file.
- Is browser TTS free?
- Yes. The voices come with the browser or the device, and speaking with them needs no API key, no account and no payment. Remote voices do use the network every time they speak. What it costs you is control: you cannot choose which voices your visitors have, and none of them is your own. The VoiceLabs API, which a page uses to speak in a cloned voice, is part of VoiceLabs Pro: US$8 a month or US$59 a year, with a 7-day trial that asks for a card up front.Pricing
VoiceLabs Pro is US$8 a month or US$59 a year, with a 7-day trial. Voice cloning and the API are included.
Microsoft, Microsoft Edge, Google and Google Chrome are trademarks of their respective owners. VoiceLabs is not affiliated with, endorsed by, or sponsored by either company. The voices a browser offers change with its version and the device it runs on.