Quick answer
TOEFL Speaking in 2026 runs eight minutes and contains eleven items in two parts. First come seven Listen and Repeat sentences, where you hear a sentence and say it back. Then comes Take an Interview: four questions from a recorded interviewer, 44 seconds of recording each, with no preparation time.
Eight minutes, eleven items, two parts
The Speaking section is the shortest part of the TOEFL. It runs eight minutes and contains eleven items. That is less time than you will spend on a single Reading module, and it is the reason so many candidates walk out of the test unsure what just happened. There is no warm-up, no long passage to settle into, and no place to hide a slow start. Eight minutes after the section opens, your entire Speaking band has been decided.
The eleven items split into two blocks that always run in the same order. The first block is Listen and Repeat: seven items where a sentence plays through your headphones and you say it back. The second block is Take an Interview: four questions asked by a pre-recorded interviewer, each answered into a 44 second recording window. There is no third block, no surprise task, and no optional section. If you know these two blocks cold, you know the entire section.
What makes this format hard is not difficulty of content. Nothing in Speaking requires academic knowledge, a specialised vocabulary, or an opinion you have to defend against a real person. What makes it hard is the pace. Every item starts on a timer you do not control, and every item ends on a timer you do not control. Your job across these eight minutes is to be ready to make sound the instant the machine asks for it, and to keep making useful sound until it stops.
This lesson maps the section end to end so nothing on test day is a surprise. The rest of the course then takes each block apart: how the scoring works, what a good repetition sounds like, how to shape 44 seconds of speech with no time to plan it, and which habits quietly hold most candidates a full band below where they should be.
The full item map
| Items | Task | What happens | Your window |
|---|---|---|---|
| 1 to 7 | Listen and Repeat | A scene picture is on screen. One sentence plays. You repeat it exactly. | A short recording window that grows with the sentence, roughly 8 seconds for the shortest and about 12 for the longest in our practice tests. |
| 8 | Interview, personal recall | The interviewer asks about something you have actually done. | 44 seconds, recording starts automatically |
| 9 | Interview, preference | The interviewer asks which of two things you prefer, and why. | 44 seconds, recording starts automatically |
| 10 | Interview, agree or disagree | A general claim is put to you and you take a side. | 44 seconds, recording starts automatically |
| 11 | Interview, policy or prediction | A broader question about effects, consequences or what should happen. | 44 seconds, recording starts automatically |
Every item you will meet, in the order it appears.
Part one: Listen and Repeat, items 1 to 7
Listen and Repeat is the quietest scoring opportunity on the test and the one candidates most often waste. A picture of a familiar place appears on screen: a university library, an advising office, a chemistry lab, a cafeteria, a train station. In our practice tests each of the sixteen Speaking sets uses one scene like this, and the seven sentences all belong to that world. Then a sentence plays once and the recording window opens. You say the sentence back.
The sentences start short and get longer. Early items are things like a one clause greeting. Later items carry two clauses and a subordinating word, which is where memory starts to fail people. The recording windows grow with the sentences, so you are never asked to say a twelve word sentence inside an eight second window, but you are also not given room to try twice.
Two things are being measured here and only two. First, did you reproduce the words accurately. Second, did you say them the way a fluent speaker would say them, with the right stress, rhythm and intonation. Nothing about opinion, content or creativity applies. This is the purest test of your ear and your mouth working together that the exam contains, which also makes it the fastest part of the section to improve. A candidate who is weak on the interview can still bank most of the available marks on these seven items inside a few weeks of daily practice.
The scoring language in our practice sets makes the ladder explicit. A perfect repetition with natural intonation and rhythm sits at the top. One minor error in pronunciation or rhythm sits one step down. Several errors with the sentence still intelligible sits in the middle. Significant errors that make you hard to understand sit near the bottom, and an unintelligible or missing response scores lowest. Notice that the gap between the top two levels is not about words at all. It is about music.
Part two: Take an Interview, items 8 to 11
The second block reframes the test as a conversation. A recorded interviewer introduces herself, explains that you have agreed to take part in a short research study, and then asks four questions about an everyday topic. In our practice sets the topics are deliberately ordinary: teamwork, outdoor activities, learning styles, leisure time, shopping habits, food, technology use, study habits, travel and so on. You are never asked to know something. You are asked to have a life and describe part of it clearly.
The four questions climb in difficulty in a fixed shape. Item 8 asks you to recall something personal, which is the most concrete and the easiest to fill with detail. Item 9 asks which of two things you prefer and why, which forces a clear choice. Item 10 puts a general claim in front of you and asks whether you agree, which needs a position and a reason. Item 11 is the broadest, asking about consequences, trade-offs or what should happen, which is where abstraction and time pressure meet.
Each answer is recorded for 44 seconds. The recording begins the moment the question audio finishes. You cannot pause it, you cannot restart it, and you cannot stop early to save time. When the 44 seconds are up the test moves on by itself. Four questions, four windows, no gaps in between that you control.
This is the part of the section that decides most bands, because it is where fluency, grammar, vocabulary and development are all visible at once. It is also where a small amount of structural training pays out enormously, because the difference between a candidate who fills 42 of the 44 seconds with organised content and one who stops at 22 seconds is usually not language ability. It is having a shape ready.
Try it now: 44 seconds, no preparation
Some students prefer to study in a library, and others prefer to study at home. Which do you prefer, and why?
Read the question once. Then press record and start talking within two seconds, exactly as the real item forces you to. Do not plan. Do not write anything down. Do not re-record because the first attempt was messy, because a messy first attempt is the honest measurement and it is the one worth listening back to.
Two things to notice about yourself afterwards, before you compare with the model. How long was the gap between the timer starting and your first word. And did you still have something to say at forty seconds, or did you finish at twenty and then repeat yourself.
What the eight minutes feel like, in order
- 1
Setup and microphone check
Before any scored item you confirm your microphone is working. Speak at the volume you actually intend to use in the test, not a cautious whisper. A check passed at whisper volume is a check that proves nothing, and a recording made at whisper volume is a recording the raters have to strain through.
- 2
Items 1 to 7, one sentence at a time
The scene appears, a sentence plays, the recording window opens, you repeat. Then the next sentence. The rhythm is fast and mechanical. Do not carry a mistake from item 3 into item 4: each item is scored on its own, and dwelling on a fluffed word is how candidates turn one lost point into three.
- 3
The interview introduction
The interviewer appears on video and sets up the research study framing. This is the only moment in the section where nothing is being recorded and nothing is being scored. Use it to sit up, unclench your jaw and take one slow breath. It is the last quiet moment you get.
- 4
Items 8 to 11, back to back
Question audio plays, recording starts, you speak for 44 seconds, the test advances. Then the next question. There is no break where you can recover from a bad answer, so the recovery has to happen inside the first two seconds of the next one.
- 5
The section ends by itself
After item 11 the section closes. There is no review screen, no chance to re-record and no confirmation that any individual answer was good. This is normal and it is not a sign that something went wrong.
Where the raw points sit
Speaking is built from 55 raw points: five points available on each of the seven Listen and Repeat items, which is 35, plus five on each of the four interview answers, which is 20. Those raw points are then converted to your reported band on the 1.0 to 6.0 scale.
Read that split again, because it changes how you should spend your practice time. Almost two thirds of the available raw points sit in the block that most candidates treat as a warm-up. Listen and Repeat is not the appetiser before the real test. It is the larger half of your Speaking score, and it is mechanically simpler to improve than interview fluency because there is exactly one right answer per item and you can hear immediately whether you produced it.
That does not make the interview optional. The four interview answers are where the ceiling is: a candidate who repeats sentences well but freezes for 44 seconds four times in a row cannot reach the top bands. But the fastest early gains almost always come from the seven repetition items, and a student with three weeks left should be spending real time there rather than only rehearsing opinions.
What you should be able to say yes to before test day
- I can describe both blocks of the Speaking section from memory, in order, with the item counts.
- I know that the interview recording starts the instant the question audio ends, with zero preparation time.
- I have practised repeating a twelve to fifteen word sentence after hearing it once, without writing it down.
- I have recorded at least one full 44 second answer and listened back to the whole thing without skipping.
- I have a first sentence I can produce for any interview question within one second of the audio ending.
- I have tested my actual microphone and headphones on the device I will use, at the volume I will use.
- I know that the section ends automatically and that there is no review screen, so I will not be waiting for one.
Check you have the format right
These three questions cover the facts that most often cause a bad surprise on test day. Get them wrong now rather than in the exam room.
How much preparation time do you get before an interview answer?
The interview gives no preparation time at all. The question audio finishes and the 44 second recording window opens in the same moment, which is why your opening line has to be automatic rather than composed.Fifteen seconds of preparation belonged to an older TOEFL Speaking design. Preparation ahead of the recording is not part of this format, and rehearsing a fifteen second planning ritual will cost you the opening of every answer.No interview item has preparation time, and the last question is not treated differently from the first three. It is broader in scope, not longer in setup.You never start the recording yourself. The test controls both the start and the stop, which is exactly why candidates who wait for a cue lose seconds.Which block carries more of the available raw points in Speaking?
Seven repetition items at five points each is 35 of the 55 raw points in the section. The four interview answers carry 20. Length of speech and weight of score are not the same thing.The interview answers are far longer to produce, but there are only four of them. Time spent speaking is not the same as points available, and assuming it is leads candidates to neglect the bigger block.They are not equal. The split is 35 raw points against 20, which is close to two thirds against one third.The topic changes between practice sets, but the item counts and the point structure do not change with it.On a Listen and Repeat item you reproduce every word correctly but say the sentence flatly, with even stress on all words and no rise or fall. What is the likely effect?
The top of the scale asks for perfect repetition with natural intonation and rhythm. Correct words with flat delivery is exactly the profile that sits one step below the top, because the words and the music are scored together.Accuracy of words is necessary but not sufficient. If word accuracy were the whole scale, there would be no way to distinguish the top two levels of the criteria used in our practice sets.Nothing is replayed. Each sentence plays once and the window opens once. There is no mechanism to repeat an item.Unintelligible means the rater cannot recover what you said. Flat but accurate speech is perfectly intelligible, it is just not native-sounding, so it lands in the middle of the scale rather than the bottom.What the rest of this course covers
The next lesson opens the scoring up: what raters are actually listening for on each block, why an answer can be grammatically clean and still score in the middle, and what separates a four from a five on the same content. After that the course splits by task. Two lessons take Listen and Repeat apart, first the core technique of hearing and holding a sentence, then the stress, rhythm and intonation that decide the top of the scale. Two more take the interview apart, first the shape of the four question progression and then a framework you can run inside 44 seconds without preparation.
The final block is habits. Fluency and pausing, which is mostly about what you do instead of filler. Paraphrasing under time pressure, for the moment the exact word will not come. Recovery moves for the answers where your mind simply empties. Microphone and room setup, which sounds trivial and quietly costs real points. And a closing lesson on the habits that cap otherwise strong candidates at a band below their level.
Work through them in order if you have time. If you have a week, do this lesson, the scoring lesson, the interview framework lesson and the recovery lesson, and spend every spare fifteen minutes recording repetitions. That sequence gets the most points back for the least time.
Common questions
Eight minutes, containing eleven items. Seven of those are Listen and Repeat sentences and four are interview questions answered in 44 second recording windows. It is the shortest of the four sections, sitting alongside Reading at 30 minutes, Listening at 29 minutes and Writing at 23 minutes in a full test of about 90 minutes.
No. There is no preparation window anywhere in the Speaking section. In Listen and Repeat the recording opens as soon as the sentence finishes playing. In Take an Interview the 44 second recording starts as soon as the question audio ends. Any planning you do has to happen in the first two or three seconds while you are already talking.
No. Every item plays once and records once, and the test advances by itself. This is why recovery technique matters more than perfection: talking through a stumble costs you very little, while stopping and restarting costs you seconds you cannot get back.
Everyday ones. Across our sixteen Speaking practice sets the topics include teamwork, outdoor activities, learning styles, leisure time, shopping habits, food and meals, technology use, study habits, health, travel and commuting. Nothing requires specialist knowledge. The difficulty is producing organised speech quickly, not finding something to say.
The section carries 55 raw points: five on each of the seven Listen and Repeat items for 35, and five on each of the four interview answers for 20. Those raw points convert to your reported band on the 1.0 to 6.0 scale, in half band steps.
In our full practice tests Speaking comes after Reading, Listening and Writing, which means you reach it with your concentration already spent. That is worth rehearsing: practising Speaking cold and fresh is easier than practising it after an hour of other work, and only one of those matches test conditions.