TOEFL Speaking scoring in 2026: 55 raw points across 7 Listen and Repeat items and 4 interview answers, what raters listen for, and what separates 4 from 5.
10 min readMonthly & Premium
Quick answer
TOEFL Speaking carries 55 raw points: five on each of the seven Listen and Repeat items and five on each of the four interview answers. Those points convert to a band from 1.0 to 6.0 in half steps. Raters weigh accuracy and delivery on repetitions, and delivery, language, development and fluency on interview answers.
Where the 55 points actually live
Most candidates study Speaking as if it were one skill with one score. It is not. It is two different scoring events wearing the same section name, and they reward different things. Understanding the split is the difference between practising the thing that will move your band and practising the thing that feels most like studying.
The section holds 55 raw points. Seven Listen and Repeat items carry five points each, which is 35. Four interview answers carry five points each, which is 20. Those raw points convert into your reported band on the 1.0 to 6.0 scale, in half band steps, aligned to CEFR levels in the same way as Reading, Listening and Writing.
Two consequences follow immediately. First, the repetition block is worth more than the interview block, by a wide margin, even though the interview takes far longer and generates far more anxiety. Second, every single item is scored independently on its own five point ladder, which means one catastrophic answer cannot sink the section. A candidate who freezes completely on item 11 and scores nothing there still keeps everything earned on the other ten items. That is worth knowing on test day, because the belief that one bad answer has ruined everything is what turns one bad answer into three.
ETS scores the real exam using a combination of automated speech scoring and trained human raters. On our practice tests, your recordings go to expert evaluation and come back with a per item score and written feedback, which is the closest thing you can get to hearing what a rater heard. In both cases the underlying question is the same: how much work does the listener have to do to understand and follow you.
The Listen and Repeat ladder
Score
What it sounds like
What usually caused it
5
Perfect repetition with natural intonation and rhythm.
You held the whole sentence in memory and delivered it as speech, not as a list of words.
4
One minor error in pronunciation or rhythm.
A single swallowed ending, one misplaced stress, or one word slightly off. Everything else landed.
3
Several errors but the sentence is still mostly intelligible.
Usually memory, not pronunciation. The back half of a long sentence drifts because you were still processing the front half.
2
Significant errors, hard to understand.
Words dropped or replaced, often with a restart in the middle. The listener has to reconstruct what you meant.
1
Unintelligible, or no response at all.
Silence, a false start that never recovered, or a recording made too quietly to hear.
The five point scale used on each of the seven repetition items in our practice sets.
What raters listen for on the repetition items
Two dimensions, weighted together. The first is accuracy: did the words you produced match the words you heard. The second is delivery: did you sound like someone speaking English or someone reciting a word list. The top of the scale requires both. This is why the most common ceiling on this block is not a pronunciation problem at all, it is a memory problem that turns into a delivery problem.
Here is the mechanism. On a short sentence you hear it, hold it whole, and say it back as one shape. On a long sentence you hear it, start repeating before you have finished processing the end, and your delivery flattens because your attention has moved from producing speech to retrieving words. The words may still all arrive. The music does not. That is the difference between a five and a four, and across seven items it is the difference between 35 raw points and 28.
The second most common loss is word endings. English marks a lot of grammar at the ends of words, and under pressure those endings are the first thing to go. A dropped plural, a dropped past tense, a swallowed final consonant. Each one is small. Each one is a point of contact where the listener has to work harder, and the scale is built out of exactly that.
What is not scored here is your accent. Having a first language that is audible in your English is not an error and does not cost you anything. What costs you is any feature that makes a listener stop and re-parse: stress on the wrong syllable of a content word, a vowel that lands on a different word than the one you meant, a rhythm so even that phrase boundaries disappear. Those are trainable and they are what the two Listen and Repeat lessons in this course drill.
The interview ladder
Score
The rubric description
The typical recording
5
Clear, well developed response with specific detail and natural delivery.
About 40 to 43 of the 44 seconds used. A position early, a concrete example, one extension, a clean close.
4
Generally clear with adequate detail, minor hesitations.
The content is there but a little thin, or a couple of visible searches for words interrupt the flow.
3
Partially developed with limited detail, noticeable pauses.
Often a 25 to 30 second answer that stopped early, or one idea repeated in three different ways.
2
Underdeveloped with vague detail, frequent pauses.
Generalities only, no example, and long gaps where the candidate was thinking rather than speaking.
1
Very limited, almost no relevant content.
A few disconnected sentences, or an answer to a question that was not asked.
The scale used on each of the four interview answers in our practice sets.
The four things audible in every interview answer
Whatever the topic, a listener forms an impression along four lines at once, and they are the four lines the rubric language points at.
Delivery. Can you be understood without effort. This covers clarity of sounds, whether your intonation moves, and whether your volume stays where the microphone can find it. Monotone is the quiet killer here. It does not cause a single identifiable error, so candidates never notice it in themselves, and it drags every answer down by roughly the same amount.
Language use. Grammar accuracy and vocabulary range together. Note the word range. Perfect grammar in a hundred word vocabulary is a mid score, not a high one. So is ambitious vocabulary you cannot control. What reads as strong is precise ordinary words used correctly, with two or three sentence patterns rather than one repeated eight times.
Development. Did you actually answer the question, and did the answer go somewhere. This is where most points are lost and where they are most easily recovered. An answer that states a position and then repeats it in different words is undeveloped even if every sentence is flawless. An answer that states a position, gives a reason, and then gives one specific instance is developed even if the English is ordinary.
Fluency. Pace and smoothness. The specific thing that hurts here is the unfilled pause: a stretch of complete silence while you search for a word. A short spoken hesitation costs far less than the same length of silence, because silence gives the listener nothing to score and reads as breakdown rather than thinking.
How to score your own recording in four minutes
1
Time it before you judge it
Look at the length first. Under 30 seconds and you have a development problem, whatever else is true. Over 44 and you were cut off, which is a planning problem. Write the number down before you listen, so your ear does not talk you out of it.
2
Count the silences over two seconds
Play it once and mark every gap where nothing at all is happening. Two or more of these in a 44 second answer is what "noticeable pauses" in the rubric describes. This one number predicts your fluency score better than anything else you can measure without training.
3
Find the position sentence and note when it arrives
The sentence that actually answers the question should be finished well inside the first fifteen seconds. If you cannot find it at all, that is a development score of three or below no matter how fluent the speech was.
4
Count concrete nouns
Names, places, numbers, times, objects. A developed answer has several. An undeveloped one has abstractions only: people, things, situations, experiences. This is a fast and surprisingly reliable proxy for the specific detail the top band asks for.
5
Count how many times you used your most frequent word
Usually it is "good", "thing", "very", "like" or "really". Four or more uses of the same filler word in 44 seconds is a visible vocabulary range problem, and it is the easiest one on this list to fix.
Why a clean answer can still score in the middle
This is the question that arrives in our inbox most often after a Speaking result. The candidate listens back, hears no grammar mistakes, no long pauses and a confident voice, and cannot understand why the score is a four rather than a five.
Usually it is one of three things. The answer was correct but generic, so it could have been given to any question on that topic and contained nothing that only this person could have said. Or the answer was well formed but flat, delivered with even stress across every word, so the listener had to supply the emphasis themselves. Or the answer used one sentence pattern the whole way through: eight sentences all built as subject, verb, object, with no subordination and no variation in length.
None of those are errors. That is exactly why they are hard to find in your own recording. They are the difference between speech that is correct and speech that is easy to listen to, and the upper half of the scale is measuring the second thing. The practical route out is not more grammar study. It is specificity, one deliberate variation in sentence shape per answer, and letting your voice move.
Test your reading of the rubric
Each of these describes a real recording. Decide what the rubric would do with it, then read why.
A candidate gives a 28 second interview answer with flawless grammar, good pronunciation and no pauses, then stops because they have said what they think. What is the most likely score?
Around 28 seconds against a 44 second window is roughly a third of the response unused, and the middle of the scale is described as partially developed with limited detail. Accuracy cannot compensate for content that was never produced.The top band asks for a well developed response with specific detail, not just accurate language. A short answer fails the development requirement no matter how clean it is.A one is for very limited content or an unintelligible recording. Twenty eight seconds of clear, relevant, accurate speech is well above that floor.Accuracy and development are separate dimensions and a strong score on one does not lift a weak score on the other. This is precisely the trade candidates hope exists and it does not.
On a Listen and Repeat item, a candidate produces every word correctly but pauses for about a second in the middle of the sentence while retrieving the second half. What does the rubric do?
The top level requires natural intonation and rhythm. A mid sentence retrieval pause is a rhythm break, so the item lands at the level below, described as one minor error in pronunciation or rhythm.If word accuracy were the whole scale there would be no way to separate the top two levels, which are distinguished by exactly this kind of delivery detail.Unintelligible means the listener cannot recover the sentence. Every word arrived correctly here, so the response is fully intelligible and sits near the top of the scale, not the bottom.Nothing is ever replayed. Each sentence plays once, the window opens once, and the section moves on.
Which of these would most improve a consistent band 4 interview performance?
The gap between four and five is described as the difference between adequate detail and specific, well developed detail. One concrete example per answer attacks exactly that gap, and it also fills time honestly, which fixes the length problem at the same time.Inserted academic phrases are the words you control least well, so they are the ones most likely to be used slightly wrongly, and that lands directly on language use. Ordinary words used precisely score better.Pace is part of the delivery and fluency impression. Speeding up degrades clarity and rarely adds usable content, so it trades a dimension you were passing for one you were not.A memorised answer collides with a question you were not expecting and produces content that does not fit the prompt, which the development dimension punishes hard. Memorise structure, never content.
A scoring self-audit you can run on any practice recording
The answer runs between 40 and 44 seconds, not 25 and not cut off mid word.
The sentence that answers the question is complete within the first fifteen seconds.
There are no silent gaps longer than about two seconds.
There is at least one concrete detail that only I could have supplied.
No single filler word appears more than three times.
At least two different sentence shapes appear, not eight copies of one shape.
My voice rises and falls, and the recording is loud enough to hear without adjusting volume.
On repetition items, the last three words of the sentence are as clear as the first three.
Monthly & Premium
10 more sections in this lesson
You have read the opening. The rest covers the method in full, with worked examples and practice questions that explain why each wrong answer is wrong.
All 45 lessons: Reading, Listening, Writing and Speaking
Real exam audio and record-yourself speaking drills
Fifty five raw points. Seven Listen and Repeat items at five points each gives 35, and four interview answers at five points each gives 20. Those raw points convert to a reported band from 1.0 to 6.0 in half band steps, aligned to CEFR levels.
No. Every item is scored on its own five point ladder, so a total blank on one interview question costs you that item and nothing else. The real damage from a bad answer comes from carrying it into the next one, which is why recovery technique is worth practising.
Having an accent is not an error and does not cost you points. What costs points is anything that makes a listener stop and work: stress on the wrong syllable, dropped word endings, or delivery so flat that phrase boundaries disappear. Those are trainable and they are what this course drills.
Specific detail and natural delivery. A four is generally clear with adequate detail and minor hesitations. A five is well developed with detail that could only have come from you, delivered without visible searching. Adding one concrete example with names, times or places is the most direct route between them.
Your recordings go to expert evaluation and come back with a score for each item plus written feedback explaining what a rater would have heard. That is the part self-study cannot replace, because the specific things that cap a band are the ones you cannot hear in your own voice.
Neither extreme. Pace is scored as part of delivery and fluency, so racing to fit content in costs you on both. Aim for the speed you use when explaining something to a friend, and buy time with specificity rather than speed. A slower answer with one real example beats a fast answer with three vague ones.