New · 150 exercises
Speak out loud, and get your delivery measured the moment you finish. 150 exercises across read aloud, listen and repeat, and interview tasks. The last two are built to the 2026 TOEFL format; read aloud is a delivery drill for pace and pausing.
Two of them mirror the real TOEFL speaking tasks question for question. The third exists because you cannot fix your pace and your pausing while you are also inventing what to say. Pick one and you go straight into it.
Get a free 1 vs 1 session with a TOEFL expert. Tell us your target score and what keeps going wrong. We will send back a study plan built around it, then track your progress as you work through it. We reply within 24 hours.
Not at the end of the exercise. After each answer, while you still remember saying it. The official practice test gives you one number at the end and no way to hear yourself back, which is the gap this fills.
Your answer transcribed word for word. On the two sentence tracks it is lined up against the target so you can see which words did not land.
Press once and hear exactly what the scorer heard. Reading your transcript and hearing the gap between what you meant and what came out is the fastest correction there is.
Not a bare number. Your words per minute next to the band that scores best here, so you know which way to move and by how much.
Your longest gap and how much of the answer was silence. Our data has this U-shaped, so we tell you when you are pausing too little as well as too much. Almost every other tool only warns you about one end.
One thing we deliberately do not do: nag you about filler words. In 59,123 graded answers on this platform, a student with one or two hesitations outscored a student with none. Telling you to eliminate them would be coaching against our own evidence.
That is not a slogan. It is what our own graded submissions say. Across every test sat on this platform, Speaking comes out lower than any other section, and the interview tasks are where it falls apart.
When we looked at what actually predicts a good speaking score in that data, the strongest single factor was not vocabulary and not grammar. It was simply how much the student managed to say in the time. That is a habit, and habits respond to drilling.
Every number below is measured from your recording. Most of the thresholds behind them were set from 59,123 graded interview answers on this platform rather than picked because they sounded reasonable. Where your pauses fall comes from published research instead, and we keep that one out of the score for exactly that reason.
Those last two are worth saying out loud, because most speaking advice on the internet tells you the opposite. We checked, and our own numbers disagree with it.
Plenty of students rehearse speaking with a general assistant, and it is genuinely useful for that. It will happily generate questions and tell you your answer sounded good. What it will not do is measure.
| A general AI chatbot | This tutor | |
|---|---|---|
| Questions | Invented on the spot, and drifts from the real format | 1,500 written to the 2026 task types, on a fixed difficulty ladder |
| Your pace | Not measured | Words per minute, against a benchmark from graded answers |
| Your pauses | Not measured | Counted, timed, and classified by where in the sentence they fell |
| A score you can compare | A different opinion every time you ask | The same scoring every run, so a retake actually tells you something |
| Progress | Nothing is measured, so there is nothing to compare between sessions | Best score per exercise, kept and tracked on your dashboard |
| Advice | General speaking tips, not tied to what your recording actually showed | Built from what our graded corpus says actually moves a score |
Four things, in the order the data says they matter. The tutor is built around them, but they are worth knowing whether or not you use it.
The median answer we grade is 80 words in a 45 second window. Answers that score well average 92, and the scores keep climbing past 110 all the way to about 130. If you stop at 30 seconds because you have run out of things to say, that is the first thing to fix, ahead of grammar and ahead of vocabulary.
At the end of a clause, fluent speakers pause about as often as learners do. The difference is all the pausing in between. Stopping after "the" or "because" leaves the listener holding an unfinished phrase. Stopping after a finished clause just sounds like thinking.
A bare opinion scores at the bottom. An opinion with a reason scores in the middle. An opinion, a reason and a specific example from your own life is what the top of the scale looks like, and it also happens to be the easiest way to fill the time honestly.
The research on this kind of practice is blunt about it. A meta-analysis of speech-recognition practice found that programmes under four weeks produced almost no lasting effect, while those running five to eight weeks produced a large one. Twenty minutes most days will beat one long session the night before, and the difference is not small.
That is the last 15 in every track, so 45 exercises in total. Each track climbs through four bands. The first three stay inside real exam parameters. Hard Mode is harder than the exam on purpose, so that test day feels like a step down rather than a step up.
Short, concrete, everyday language. Build the habit of speaking without stopping.
Longer sentences and less familiar topics. This is where pace starts to matter.
Real exam length and topics, on the real clock. This is where it stops feeling generous.
Past what the exam uses, on purpose. Longer sentences, a faster prompt voice, more abstract questions, and no thinking time at all.
Reading, Listening and Writing each have their own task-by-task drills, built the same way: one task type at a time, levels that climb to the real exam and then past it, and every answer explained. Level 1 of every task type is free.
Word count came out as the strongest predictor in the whole set, at a correlation of 0.49 with the score the answer received. The curve is steep and it is close to monotonic. Answers of 0 to 30 words average 0.99 out of 5. Answers of 90 to 110 words average 2.82. Answers of 110 to 130 average 2.97. Answers of 130 words and up average 3.20. Nothing else in the data moves a score that far. If you stop talking at second 30 because you have run out of things to say, that is the first thing to fix, ahead of grammar and ahead of vocabulary, because everything else you might improve is being applied to half an answer.
Pace follows word count closely, at a correlation of 0.35. High scoring answers run at about 137 words per minute. Low scoring answers run at about 112. The gap is 25 words a minute, which over a 45 second window is roughly 19 words, and 19 words is often the difference between a claim with an example attached and a claim left hanging. Note what 137 words per minute actually means: in a 45 second answer it produces about 103 words, which lands inside the target band our own data says scores best. It is a brisk conversational speed, not a race.
Silence is the finding that surprises most students, because it does not behave the way the advice on the internet says it does. It is U-shaped, not linear. Answers that were 15 to 25 percent silence scored best, at 2.83. Answers with almost no silence at all, 0 to 15 percent, scored lower at 2.67. Answers that were 45 percent silence or more collapsed to 1.78. Speaking without a break does not read as fluent. It reads as rushed, and it usually means you are reciting rather than thinking. The goal is not to eliminate pausing. It is to move your pauses to the ends of clauses, where a listener expects them, and out of the middle of phrases, where they leave the listener holding an unfinished sentence.
Filler words barely register. The correlation between hesitation sounds and the score is -0.10, which is close to nothing. Answers with one or two fillers averaged 2.60, very slightly higher than answers with none at all, which averaged 2.57. Only at six or more does anything measurable start to cost you. A student who spends a week trying to stamp out every um is optimizing a variable with almost no weight while the one variable that carries half the signal, how much you said, goes untouched.
The failure rate tells you where to spend your time. Fifty-five percent of interview answers in the corpus score 2 out of 5 or below, and the most common reason is simply not saying enough. On the repetition task, forty percent of attempts score 2 or below, and almost all of that is the difficulty of holding a whole sentence in memory rather than anything to do with pronunciation. Those two numbers are why this product exists, and why its report screen is built around length, pace and pausing instead of a single band.
The 2026 TOEFL iBT Speaking section runs about 8 minutes and contains 11 items worth 55 raw points: 7 Listen and Repeat items, then 4 interview questions at 45 seconds each. Two of the three tracks here map onto those task families one for one. The third does not, and it is deliberate.
Listen and Repeat is 50 exercises of 10 items each. You hear a sentence once and repeat it back, with the text hidden until after your attempt, exactly as the real task works. Sentences start at 6 to 10 words in the early exercises and climb to 17 to 26 words at the top, with the recording window moving from 8 seconds to 15 as the sentences grow. What the drill actually trains is working memory under a single hearing, which is where the failures are. Students who assume their repetition score is an accent problem usually find, once they see which words dropped out, that they lost the back half of a sentence they never fully held.
Interview is 50 exercises of 10 questions each, always on a 45 second clock, because shortening the answer window would make the drill easier rather than harder and the whole problem this product exists to fix is students not filling the window. The difficulty ladder moves thinking time instead. Preparation runs 10 seconds in the first eight exercises, 5 seconds through the middle, 3 seconds at test pace, and 0 seconds at the top, which is what the real section gives you anyway. Each exercise carries a target word range taken straight off the measured curve: 70 to 100 words while you are warming up, 90 to 130 once you are building, and 110 to 140 in the hardest band.
Read Aloud is 50 exercises and it is not a TOEFL task at all. A sentence appears on screen and you read it out loud, with windows of 13 to 22 seconds depending on length, and sentences climbing from 8 to 14 words up to 20 to 34. It is here because of what the data above says. Word count and pace are the two things that carry the score, and a student cannot fix pace and pausing while simultaneously inventing an argument against a running clock. Read Aloud removes the composition load so the delivery habit can be built on its own. Interview then puts the load back. If you only have time for one track before a test date, do Interview, since that is where the marks are lost. If you have three weeks, start on Read Aloud for the first few days and you will arrive at Interview already speaking at a usable rate.
Every interview exercise rotates through all four question shapes so you are never drilled on just one: personal recall, preference, agree or disagree, and prediction. That rotation is a design decision backed by the corpus. Average scores across the four types run from 2.21 to 2.43, a spread small enough that no single type is anyone's real weakness. Students who believe they are bad at opinion questions are usually just as short on personal recall questions. The problem is length, and length is type independent.
A full answer has the same skeleton whatever the shape of the question. Say your position in one sentence so the listener knows where you are going. Give one reason for it. Then give a specific example from your own life, with a name, a place, a number or a date in it. A bare opinion sits at the bottom of the scale. An opinion with a reason sits in the middle. An opinion, a reason and a concrete example is what the top looks like, and it is also the easiest honest way to fill 45 seconds, because a real example has detail in it and detail is what generates words.
If you want a number to aim at, 103 words is what 45 seconds at 137 words per minute produces, and that sits comfortably inside the 90 to 130 band the data says scores best. Count it once on a written answer so you know what 100 words feels like out loud. Most students are shocked at how little it is. Roughly four sentences of normal length.
Each track climbs through four named bands. Exercises 1 to 8 are Warm Up: short, concrete, everyday language, built to establish the habit of speaking without stopping. Exercises 9 to 20 are Building: longer sentences and less familiar topics, which is where pace starts to matter. Exercises 21 to 35 are Test Pace: real exam length and topics on the real clock. Exercises 36 to 50 are Hard Mode, which is past what the exam uses, on purpose. That is 15 Hard Mode exercises in every track, 45 in total.
Hard Mode raises three levers at once. Sentence length goes beyond the range the real section uses. The prompt voice speeds up, by 12 percent on Listen and Repeat and 10 percent on Interview, which on a repetition task is the single biggest difficulty lever there is. And thinking time drops to zero. The interview questions get more abstract at the same time, and the target answer length rises to 110 to 140 words. The point is not to be unfair. The point is that if your practice ceiling is the exam, then on test day you are operating at your ceiling with adrenaline on top. If your practice ceiling is above the exam, test day is a step down.
Scoring on each attempt is a percentage, and the star cuts are fixed and public so a retake means something: 85 and up is three stars, 70 is two, 50 is one. Only your best attempt at an exercise is kept, so a bad retake can never cost you anything. Progress accumulates as experience points across eight levels, from Getting Started up to Command, weighted by both exercise difficulty and score so that repeating exercise 1 forever is worth far less than clearing exercise 40 once.
Measured, on every attempt: how much you said, your pace in words per minute against the 137 that strong answers average, how much of the answer was silence, how long your longest single stop ran, how long you took to start speaking, and on the two sentence tracks, how much of the sentence actually landed, word by word. Where your pauses fell is also reported, split by whether they came at a clause boundary or inside a phrase, since published research is clear that fluent speakers pause at clause ends about as often as learners do and that the extra pausing a learner does is almost all mid sentence. That one is shown to you but kept out of the score, because the thresholds behind it come from research rather than from our own measured corpus, and we hold scored components to the corpus.
Not measured: pronunciation and accent. Doing that honestly needs sound by sound analysis that our speech recognition does not produce, so any accent score we displayed would be invented. On the two sentence tracks we do check which words the recognition heard, so speech that is very hard to make out can still cost you there, but that is a word match and not a verdict on how you sound.
Also not treated as faults: a handful of filler words, and silence itself. The data above is the reason. One or two hesitation sounds score marginally higher than none, and 15 to 25 percent silence scores better than near continuous speech. Most speaking advice online tells you the opposite on both counts. We checked our own numbers and they disagree, so we score what the numbers support and we tell you which parts we are leaving alone.
Give it weeks rather than days. The research on speech recognition based speaking practice is blunt about the timeline: programs running under four weeks produced almost no lasting effect, while programs running five to eight weeks produced a large one. Twenty minutes on most days beats one long session the night before by a margin that is not close. A single exercise is ten questions, which is about two minutes on Read Aloud and about ten on Interview once you include the clock, so a twenty minute session is two or three exercises.
Do not grind one exercise. Lambert, Kormos and Minn (2017, Studies in Second Language Acquisition) had 32 learners perform three task types six times each and found speech rate gains were largest across the first three performances, continued to about the fifth, then flattened. The learners themselves judged three to four repetitions optimal. Shiki and colleagues report the same plateau at about five for repeating the same script. Retakes here are unlimited and always will be, but once you have run one exercise five times the next exercise is worth more to you than a sixth attempt at this one, and the report screen says so.
Spread the repetitions out rather than stacking them. Suzuki and Hanzawa (2022, Studies in Second Language Acquisition) compared massed against spaced repetition and found massed repetition cut pausing the most in the moment, but produced slower articulation and more verbatim repetition, and no scheduling effect survived to a delayed post test a week later. In practice: three exercises today and three tomorrow will serve you better than six in one sitting, even though six in one sitting will feel more productive while you are doing it.
A workable four week shape looks like this. Week one, Read Aloud exercises 1 to 8 plus Interview 1 to 4, so the delivery habit starts forming while you meet the question types. Week two, Listen and Repeat 1 to 12 and Interview 5 to 12, and start watching your word count rather than your star rating. Week three, Test Pace on all three tracks, exercises 21 to 35, which is where the prep time drops to 3 seconds and the sentences reach real exam length. Week four, Hard Mode on whichever track your reports are weakest on, and one full timed speaking test at the end of the week to see whether it moved.
They do different jobs and you want both. A speaking practice test is one timed run through the whole section, sat once, producing a band and an expert evaluation of what you produced. It is a measurement. These exercises are the drilling: ten items on a single task type, repeatable as often as you like, with only your best attempt kept and a report about delivery rather than a band. A measurement tells you where you stand. A drill is what changes it.
The loop that works is straightforward. Sit a speaking practice test to find out which task family is costing you, drill that family here for two or three weeks against your own word count and pace numbers, then sit a different test and compare. If your interview word count has moved from the 60s into the 100s and your pace has come up toward 137, the band will usually have followed. If you want the underlying theory instead of the reps, the TOEFL Speaking course cover how the section is scored, the interview framework, and what to do when you lose your thread halfway through an answer.
These are the speaking drills. Reading, Listening and Writing have their own, built the same way and covering all ten of the task types the 2026 test uses: the full set of TOEFL practice exercises runs one task type at a time, twenty levels each, with every answer explained. Level 1 of every task type is free, the same as exercise 1 here.
One hundred and fifty exercises of ten questions each, 1,500 questions in total, split across three tracks of 50: Read Aloud, Listen and Repeat, and Interview. You answer every one out loud in the browser and get a delivery report the moment the exercise ends, covering how much you said, your pace in words per minute, how much of the answer was silence, where your pauses fell, and on the two sentence tracks how much of the sentence came through word by word.
A practice test is one timed run through the whole Speaking section and it produces a band with an expert evaluation. This is drilling. Each exercise is ten questions on a single task type, you can retake it as often as you like, only your best attempt is kept, and the report is about how you delivered rather than what band you would get. Most students get the most out of doing both: sit a test to find the weak spot, drill it here, then sit another test to see whether it moved.
Exercise 1 of every track is free, which is 30 questions across Read Aloud, Listen and Repeat and Interview, and it needs no signup. That is enough to see your own word count, pace and silence numbers on all three task types before you decide anything.
1,500. Every exercise is 10 questions, there are 50 exercises per track, and there are three tracks. On the Interview track the questions rotate through all four shapes the section uses, which are personal recall, preference, agree or disagree, and prediction, so no exercise leaves you drilling a single question type.
No, and we will not claim to. Scoring pronunciation properly needs sound by sound analysis that our speech recognition does not produce, so anything we told you about your accent would be guesswork. What we measure is delivery: speaking rate, the length and position of your pauses, how much of a sentence you reproduced, and how much you said in the time. On the Read Aloud and Listen and Repeat tracks we do check which words the recognition heard, so speech that is very unclear can cost you there, but that is a word match rather than a verdict on how you sound.
Read Aloud runs about two minutes for ten items, since each recording window is 13 to 22 seconds. Listen and Repeat is a little quicker per item, at 8 to 15 seconds a window, plus the time to hear each sentence. Interview is the long one at roughly ten minutes, because every answer is a full 45 seconds plus thinking time. A twenty minute session is two or three exercises.
Exercises 36 to 50 in every track, 45 exercises in all, set deliberately beyond the real exam. Sentences run longer than the section uses, the prompt voice speeds up by 12 percent on Listen and Repeat and 10 percent on Interview, the interview questions get more abstract, the target answer length rises to 110 to 140 words, and thinking time drops to zero, which is what the real section gives you anyway. The aim is that test day feels like a step down rather than a step up.
As many times as you like, and only your best score is kept, so a retake can never cost you anything. One thing worth knowing from the research on repeated speaking practice: gains are largest across the first three attempts, continue to about the fifth, then flatten. Once you have run an exercise five times you will learn more from the next exercise than from a sixth try at this one.
Plan on four at minimum and prefer five to eight. Studies of speech recognition based speaking practice found that programs shorter than four weeks produced almost no lasting effect while programs of five to eight weeks produced a large one. Twenty minutes on most days is a better use of the same total time than one long session, and spacing the repetitions across days rather than stacking them in one sitting is what survives to a delayed test.
Yes. It records through the browser, so there is nothing to install. You will need to allow microphone access, and headphones help on the Listen and Repeat track so the prompt audio does not bleed into your recording.
The first exercise in each track is free and needs no signup. That is ten questions: about two minutes on Read Aloud, or ten on Interview if you want the full picture.
Prefer to read first? The TOEFL Speaking course covers how the section is scored, the interview framework, and what to do when you lose your thread mid answer.