How to Use Text-to-Speech in English Classes
When text-to-speech helps ESL teaching and when it does not, with practical classroom patterns for using neural voices for listening, pronunciation, and differentiation.
Text-to-speech (TTS) has moved from robotic menus to surprisingly natural neural voices. For ESL teachers, this is genuinely useful — but only if you use it for the right things. This guide explains where TTS adds value in class and where it is still better to use a real recording.
What neural TTS is good at
Modern neural TTS (the kind that runs in your browser or on a server with a voice model) is good at:
- Reading short, scripted dialogues at adjustable speed.
- Producing consistent pronunciation for repeated drills — every learner hears the same model answer.
- Generating audio on demand for any text, in seconds, without booking a recording studio.
- Customising voice and speed to match the level of your learners.
- Running offline, in a browser tab, without uploading student data.
Where it is still limited
TTS is not yet a replacement for:
- Real spontaneous speech with false starts, fillers, and natural rhythm.
- Accents outside the supported voices — if you teach a specific regional accent, check the voices first.
- Pronunciation models for difficult connected speech — TTS still occasionally flattens intonation in ways that a teacher can model more accurately.
A practical rule of thumb: use TTS for receptive tasks (listening, dictation, gap-fill) and model pronunciation for short sentences. For fluency work with learners, prefer a real voice — yours.
Five classroom patterns that work
1. Slow-fast pair
Generate the same sentence twice — once at 0.7× and once at 1.0×. Students listen to the slow version first to catch the meaning, then to the natural version to internalise the rhythm.
2. Shadow reading
Play a sentence at natural speed. Students repeat immediately, copying the pronunciation and intonation as closely as possible. This works best with shorter sentences and clearly-paced voices.
3. Dictation chains
Read a short paragraph at 0.85×. Students write what they hear, then compare with a partner before you reveal the text. Dictation is one of the strongest listening activities, and TTS gives you a clean source every time.
4. Voice A / Voice B dialogues
For dialogues, assign a different voice to each speaker. This makes conversations much easier to follow than a single voice reading "Speaker 1 said... Speaker 2 said..." Switching voices tells the listener who is talking without any visual cue.
5. Differentiation by speed
In the same class, advanced learners can listen at 1.0× while beginners listen at 0.75×. With TTS, generating both versions takes seconds, and you can hand out two MP3 files from the same lesson.
A quick workflow for a TTS-based lesson
- Pick or write the dialogue you want to use.
- Open a browser-based TTS tool.
- Choose a clear voice and a slow-to-medium speed (0.85× is a good default for ESL listening).
- Generate the audio, listen once to check pronunciation.
- Download the MP3 and drop it into your slide deck, learning platform, or just play it from the browser.
For a step-by-step guide on how to design the listening task itself, see how to create listening exercises for ESL students.
What to listen for when evaluating a TTS voice
Before using a TTS voice in front of a class, check:
- Stress and intonation. Are content words clearly stressed? Does pitch fall on statements and rise on yes/no questions?
- Connected speech. Does the voice handle contractions naturally? Are word endings clear (especially -ed and -s)?
- Pacing. Is the default speed too fast for your learners? Can you slow it down without making it sound unnatural?
If the answer to any of these is "no," choose a different voice or slow the audio down.
Privacy and data
If you use a TTS tool that runs in the browser (no upload to a server), no student text ever leaves the device. This is the right choice for any classroom context where student writing might contain personal information.
If a TTS service uploads text to a remote API, check:
- whether the text is stored or only processed in memory
- whether the provider trains new models on user input
- whether you need to anonymise student names before pasting
Bringing it together
TTS is a force multiplier for teachers who already know what to do with audio. It does not replace lesson planning — it removes the friction from producing the audio you have already planned to use. A few minutes of setup can save hours of searching for the right recording, and you can customise the result to the exact level of the learners in front of you.
For a catalogue of pre-written dialogues organised by level, see A1 listening exercises, A2 listening exercises, and B1 listening exercises.