Text to Speech: Listen to Anything, Free & Private
Your device already contains a capable narrator. How browser text-to-speech works, the listening speeds that actually stick, proofreading by ear, accessibility uses — and the honest limits nobody puts on the label.
Reading takes your eyes. Listening only borrows your ears.
There's a growing pile most of us carry: the article saved on Tuesday, the report you'll "get to", documentation you genuinely need, the newsletter you like but never open. The bottleneck isn't interest — it's that reading demands your eyes and your posture and a block of undisturbed attention. Listening demands none of those. Turn the pile into audio and it drains itself during the commute, the dishes, the walk.
The surprise is that you don't need a subscription for this. Every modern operating system ships speech synthesis — the same engines behind screen readers and navigation prompts — and browsers expose them to web pages through the Web Speech API. Our free text-to-speech reader is a thin, honest layer over that machinery: paste anything, pick from the voices your device provides, set speed and pitch, press listen. Because synthesis happens locally, three good things follow automatically — it's free without a catch, it works offline with installed voices, and your text never leaves the tab. Paste a confidential draft or a medical letter without a second thought; there is no server to trust because no server is involved.
The reader adds the handful of touches that make the raw API pleasant: a word count with estimated listening time at your chosen speed, sentence-aware chunking so book-length pastes don't stall (a notorious engine quirk), and pause/resume that behaves. For PDFs specifically — page navigation, word-by-word highlighting — our dedicated PDF Read Aloud app does the same thing with document superpowers.
Where the voices come from (and why your list differs)
Click the voice dropdown and you might see eight voices or forty — and your friend's laptop shows a different set entirely. That's because the browser doesn't ship voices; it inventories them. Windows contributes its narrator voices, macOS its famously large multilingual collection, phones their own; some browsers add network-backed voices on top. Each voice carries a language tag, which matters more than accent aesthetics: a text's language should match the voice's, or you'll hear English phonetics gamely mangling French.
Quality varies in one important dimension: age of the voice technology. Older "robotic" voices synthesize phoneme-by-phoneme; newer neural voices — increasingly the OS defaults — model whole phrases and land startlingly close to human narration. If your list sounds dated, updating the OS or installing additional voices in system accessibility settings upgrades every app that uses them, this reader included. It's the rare software improvement you make once at the OS level and inherit everywhere.
Two knobs shape any voice further. Rate is the workhorse — more on it below. Pitch is subtler: tiny adjustments (0.9–1.1) can make an artificial voice sit better in your ear, while extreme values are mostly comedy. When a voice fatigues you on long listens, changing the voice beats torturing the pitch slider.
Listening speed: the skill worth training
Here's the arithmetic that makes speed worth caring about. Average silent reading runs 200–250 words per minute; comfortable speech lands near 150–180. At 1.0×, listening is slower than reading — you're trading speed for hands-free. But comprehension research and a million audiobook listeners agree the trade improves fast: most people acclimate to 1.2–1.6× within days, which pushes effective intake to 220–290 wpm — now faster than typical reading, still hands-free. The reader's time estimate updates with the rate slider, so a 4,000-word report visibly drops from 25 minutes to 16 as you nudge the speed.
Training it is unglamorous: start at 1.0× for a session, raise by 0.1 whenever the voice starts feeling slow, stop where comprehension dips, and park one notch below. Two caveats keep it honest. Dense technical material deserves a slower gear than newsletters — speed is per-content, not per-person. And there's a difference between hearing and absorbing: if you reach the end of a section and can't summarize it, the speed was vanity. The goal is the fastest rate at which you'd pass a quiz, not the fastest rate you can endure.
The proofreading trick writers swear by
The second-biggest use of text-to-speech has nothing to do with saving time — it's catching mistakes. When you proofread your own writing, you don't read what's on the page; you read what you meant, because your brain autocompletes its own sentences. A synthetic voice has no such loyalty. It reads the words that exist: the "the the" your eyes vaulted over, the sentence that never got its verb, the word "form" where you meant "from" — spellcheck-proof and eye-proof, but glaring to the ear.
The workflow is almost embarrassingly simple: paste your draft, drop the speed to 0.9–1.0× (slower is better for this job), and follow along with the cursor. Rhythm problems announce themselves too — the sentence you have to hear twice is the sentence your reader will have to read twice, and the paragraph that sounds breathless needs a period somewhere. Writers who adopt this habit for anything that matters — cover letters, proposals, publish-worthy posts — rarely go back to eyes-only proofing. It pairs naturally with our text utilities for the mechanical cleanup (case, duplicates, counts) before the listen.
Students discover a cousin of the same effect: hearing notes read aloud while reviewing them visually engages a second memory channel, and awkward-but-effective, playing your own summaries at 1.4× is a legitimately efficient revision pass before an exam.
Accessibility: where this stops being a convenience
For a large group of people, text-to-speech isn't a productivity hack — it's the difference between accessing text and not. Dyslexic readers consistently report better comprehension and far less fatigue listening than decoding; many use audio as the primary channel with the text visible alongside. Anyone with low vision, eye strain from screen-heavy work, migraines, or age-related sight changes gets the same door opened. And unlike full screen-reader software — powerful, but with a real learning curve — a paste-and-listen page requires no training at all: it meets people at the simplest possible interface.
The privacy property matters double here. Accessibility needs often travel with sensitive documents — medical letters, legal notices, financial statements. A reader that synthesizes locally means those documents are never uploaded to anyone's "free" service in exchange for the reading. If you support a family member who struggles with print, bookmarking a private, zero-setup reader on their machine is a genuinely useful five-minute favor.
Language learners round out the audience: hearing a native-language voice pronounce text while reading along tightens the sound-to-spelling mapping, and the speed slider lets you slow a language you're acquiring to a rate your ear can parse — 0.7× French today, 0.9× next month, real-time eventually.
Honest limits (and the workarounds)
No MP3 download. Browsers deliberately don't hand pages the synthesized audio stream, so a "save as audio" button can't exist here. The practical workaround for occasional needs: play the speech in a tab and capture it with our screen recorder using tab-audio capture — you get a WebM with the narration. Regular audio-file production is dedicated-TTS-service territory.
Voices are device-bound. Your listening setup doesn't follow you between machines; a work laptop may have a worse voice list than your phone. The fix is upstream — install better voices at the OS level on the machines you use most.
Pronunciation quirks. Engines occasionally mangle names, acronyms and domain jargon. For a one-off listen it's charming; for proofreading it's irrelevant; only for produced audio does it matter, and that's again dedicated-tool territory with pronunciation lexicons.
Long-text stalls — the classic Web Speech bug where the engine goes silent mid-chapter — are handled for you here: the reader splits input into sentence-sized utterances and queues them, so a 50,000-character paste plays straight through. It's the kind of invisible fix you only notice on tools that lack it.
Building a listening habit that actually sticks
Tools don't create habits; defaults do. The people who actually convert their reading backlog into listening time share a pattern worth copying. They pair listening with a fixed activity — the commute, the dishes, the dog — so the trigger is automatic rather than willed. They batch: once a week, collect the articles worth hearing into one paste session instead of deciding per-article. They start sessions at the speed they ended the last one (the reader keeps your settings while the tab is open), so the 1.4× that felt fast in March is invisible by May. And they accept that some material refuses the format — anything with dense tables, code, or heavy diagrams wants eyes — without letting those exceptions kill the habit for the 80% of prose that listens beautifully.
One more compounding trick: listening pairs naturally with the other direction of the pipeline. Draft by voice memo on a walk, transcribe, clean the text up with the text utilities, then proof it by ear at 0.9× before sending. Your ears catch what your eyes forgive, in both directions — and a workflow where text is heard at least once before it ships simply produces better writing.
A note on voices and the neural leap
If you tried text-to-speech years ago and bounced off the robot voice, it's genuinely worth a second audition. The last few OS generations shipped neural voices — synthesis models that learned intonation from human speech rather than assembling phonemes — and the difference is not incremental. Sentences breathe, questions rise, lists get rhythm. On a recent Windows, macOS or phone, the default voice is often good enough that first-time listeners ask which podcast app is playing. The reader inherits whatever your device offers, so the single best upgrade to your listening experience is a system update or a five-minute visit to your OS's speech settings to install a newer voice — after which every session here sounds better, at the same price of nothing.
Text isn't going anywhere, and neither is the backlog — but the hours your ears sit idle every day are a genuine, renewable resource. Point them at the pile. A voice you already own, a page that keeps your text private, and a speed slider you'll outgrow monthly: that's the whole toolkit, and it costs exactly nothing to start using it today.
Start with one article today — the one that's been open in a tab for a week. Paste it, press listen, and let the backlog begin to drain while your hands do something else entirely.
What listens well, and what doesn't
Not every document survives the trip from page to ear, and knowing which is which saves a lot of frustration. The deciding factor is how much the text depends on spatial layout. Prose is linear already, so it converts almost perfectly — essays, news, fiction, long-form reporting, documentation written in paragraphs, and email threads all listen beautifully. Anything where meaning lives in the arrangement rather than the sentences fights the format hard.
The clearest failures are tables, where a synthesiser reads cells left to right and the relationships between columns evaporate; code, where indentation carries logic that speech simply cannot express; and anything heavy with figures, footnotes or cross-references, because you cannot glance back without losing your place. Mathematical notation is the worst case of all — a formula that takes two seconds to read takes twenty painful seconds to hear, and still leaves you unsure of the grouping.
There is a useful middle category worth handling deliberately. Academic papers, legal contracts and technical specs are mostly prose interrupted by structure. The practical approach is to listen to them for shape — what the argument is, where it goes, which sections matter — and then read only those sections with your eyes. That is a genuinely faster workflow than reading the whole thing linearly, because it lets you skip the 70% of a paper that turns out not to be relevant to you without having had to read it first to find out.
A last piece of practical advice: strip the furniture before you paste. Navigation menus, cookie notices, "share this" prompts and author bios all get read aloud with the same seriousness as the article, and they break the flow badly. Copying from a browser's reader view, or running the text through the text utilities first to collapse the whitespace and drop the junk lines, takes ten seconds and noticeably improves every minute of listening that follows.
Frequently Asked Questions
Is this really free, and what's the catch?
Free because the synthesis engine is already on your device — the page just drives it. No character limits, no accounts, no upload. The honest "catch" is the limits above: no MP3 export and device-dependent voices.
Why do I see different voices on my phone vs laptop?
Voices belong to the operating system, not the page. Each device inventories its own; installing extra voices in OS accessibility settings expands the list everywhere.
What speed should I listen at?
Start at 1.0×, drift up to the 1.2–1.6× band as your ear adapts, slow down for dense material — and for proofreading your own writing, stay at or below 1.0×.
Can it read PDFs and web articles?
Anything you can paste, it can read. For PDFs with page handling and word highlighting, use our PDF Read Aloud app — same private, in-browser approach with document features.
Is my text private?
Yes — synthesis runs locally through the Web Speech API. The text exists only in your tab and disappears on refresh; with locally installed voices, it even works offline.