The TOEFL Listen and Repeat task gives you seven sentences, each played once. The chunking technique that holds a long sentence and delivers it as speech.
11 min readMonthly & Premium
Quick answer
Listen and Repeat is the first seven items of TOEFL Speaking. A sentence plays once against a scene picture and you say it back into a short recording window. You are scored on word accuracy and on delivery together, so the technique is to hold the sentence as two or three meaning chunks, not as a list of words.
Seven sentences, one chance each
Items 1 to 7 of the Speaking section are repetitions. A picture of an everyday place sits on screen, a sentence plays through your headphones once, and the recording window opens. You say the sentence back. Then the next one. The whole block is over in a couple of minutes and it carries 35 of the 55 raw points in the section.
The sentences get longer as the block goes on. Early items are one short clause. Later items carry two clauses joined by a word like "though" or "if", and by item 7 you are holding something in the region of fifteen words. The recording windows grow with them, from roughly eight seconds for the shortest sentences up to about twelve for the longest in our practice tests, so the window is never the constraint. Your memory is.
Here is what almost nobody realises until they see their scores: this block is not a pronunciation test. Candidates with excellent pronunciation routinely score three out of five on items 6 and 7, and the reason is never the sounds. It is that they started speaking before the sentence had finished arriving in memory, so the second half came out as fragments. The technique in this lesson is a memory technique that happens to produce good delivery, not a delivery technique that happens to need memory.
Hear the length ladder
The third sentence is the hardest to repeat. What structural feature is doing most of that work?
The sentence opens with a condition ("if the book you want is already on loan") and then delivers a result. You have to store both halves before you can start, which is exactly the load that makes the later items in the block harder.Every word in the third sentence is ordinary: book, loan, reserve, copy, counter. Vocabulary difficulty is not what changes across the ladder. Structure is.The sentences play at the same natural pace. What changes is how much you have to hold, not how fast it arrives.The sounds are the same sounds used in the shorter sentences. Nothing new is introduced phonetically, which is the point: the difficulty is memory load.
Show transcript
Here are three sentences from a campus scene. Listen to how each one gets longer. One. The reading room closes at nine. Two. Please return your books to the front desk. Three. If the book you want is already on loan, you can reserve a copy at the counter.
The technique: chunk, do not list
Working memory does not store words. It stores meaning units. If you try to hold "if the book you want is already on loan you can reserve a copy at the counter" as sixteen separate items, you will lose about half of them, because that is simply more than short term memory holds. If you hold it as two units, a condition and an instruction, you will keep all of it, because two is trivially inside your capacity.
So the entire technique is this: while the sentence is playing, do not try to catch words. Catch the shape. Ask yourself, in the background, how many pieces is this. Almost every sentence in the block is one, two or three pieces:
One piece: "The reading room closes at nine." Two pieces: "Please keep your voice down / when you are near the study carrels." Three pieces: "Food is not allowed near the shelves, / though water bottles are fine, / so please leave drinks at the entrance."
When the audio stops, you replay the shape, not the words. Piece one comes out and pulls its own words with it. Piece two follows. This is why fluent speakers can repeat long sentences effortlessly in a language they know well and cannot do it at all in a language they do not: the chunking depends on understanding, not on hearing.
Which leads to the most useful and least obvious rule in this lesson. Understand the sentence and the repetition takes care of itself. If you catch the meaning, your mouth can reconstruct the words. If you catch only sound, you are holding noise and it will decay before you finish speaking. Candidates who train by trying to hear harder plateau quickly. Candidates who train by making sure they understood first improve fast.
The four beats of a clean repetition
1
Beat one: listen for meaning, not for words
While the audio plays, your job is comprehension, not transcription. Picture the thing being described. A picture is a chunk, and a chunk survives in memory far longer than a string of syllables. If you find yourself silently repeating words as they arrive, stop: that habit competes with listening and it is why the ends of sentences disappear.
2
Beat two: wait for the full stop
Do not start until the sentence has actually finished. The instinct to begin early feels like it saves time, and the window is generous enough that it saves you nothing. Starting early guarantees that you are still receiving while you are producing, and that is the exact condition under which the back half of the sentence turns to fragments.
3
Beat three: start with the first chunk at full commitment
Say the opening chunk at normal speaking volume and normal speed, as a phrase, with the stress where it belongs. Do not test the water with a tentative first word. A confident opening chunk sets the rhythm for the whole sentence, and rhythm is half of what separates the top two score levels.
4
Beat four: let the sentence land
Finish the last word properly and let your pitch fall. Most lost points at the end of a repetition are not wrong words, they are trailing off: the volume drops, the final consonant is swallowed, the pitch stays flat as if you were going to continue. The sentence has to sound finished, because a sentence that sounds unfinished reads as an incomplete response.
Record one: a two chunk sentence
Listen to the model sentence below, then record yourself repeating it exactly: "Please keep your voice down when you are near the study carrels."
Do not start speaking while the sentence is still playing. In the first second after it ends, ask yourself one question: how many pieces was that. The answer here is two, a request and a condition. Then say piece one, "please keep your voice down", as one connected phrase at normal volume, and let piece two follow straight after it without a gap. Total elapsed time from the end of the audio to the end of your repetition should be under four seconds.
0:10
Model answer
Listen to the model against your own recording with three specific things in mind.
The phrase boundary. The model has one small break, after "down". That is the join between the two chunks. If your recording has breaks in other places, or no break at all, the rhythm is not matching the meaning and that is what a rater hears as unnatural.
The stressed words. "Voice", "down" and "carrels" carry the weight. The function words around them ("please", "your", "when you are near the") are compressed and fast. English rhythm is built on this contrast, and giving every word equal time is the most common single reason a word-perfect repetition scores four instead of five.
The ending. The final syllable of "carrels" is fully pronounced and the pitch falls. If your version trails off, the sentence sounds abandoned rather than finished, and that costs a point that has nothing to do with your English.
The same sentence, three ways
Versions one and two both contain every correct word. Why would both score below version three?
The scale rewards perfect repetition with natural intonation and rhythm. Version one gives every word equal weight and equal spacing, so no phrase emerges. Version two removes the boundaries entirely. Both are accurate and both are unnatural, which is exactly the profile that lands one level below the top.Volume is not what distinguishes these three. All three are audible. The distinguishing feature is how the words are grouped in time.All three versions contain the identical nine words. Nothing has been substituted, which is precisely why this comparison isolates delivery from accuracy.Length of response against the window is not the issue here. All three fit comfortably. The difference is entirely in the shaping.
Which delivery problem is more common among candidates who score in the middle of this block?
The word by word pattern is what memory strain produces. When you are retrieving each word as you go, the retrieval time inserts itself between the words and the sentence comes out evenly spaced. Rushing is rarer because it requires already having the whole sentence available.Rushing happens, but it usually means the candidate had the whole sentence in memory and simply got nervous. That is a smaller and easier problem than the retrieval spacing that produces version one.Mid scoring responses very often have every word correct. If word errors were the only cause of mid scores, this whole lesson would be about vocabulary rather than about chunking.On short sentences neither problem shows up much, because memory is not strained. Both appear on the longer items later in the block, and the word by word pattern dominates.
Show transcript
One sentence, three deliveries. Version one, word by word. Books. Can. Be. Borrowed. For. Up. To. Three. Weeks. Version two, rushed together with no phrasing at all. Bookscanbeborrowedforuptothreeweeks. Version three, the way a person actually says it. Books can be borrowed for up to three weeks.
Match your technique to the sentence length
Where you are
Typical sentence
What to focus on
Items 1 and 2
One short clause, five to seven words.
Do not overthink these. Full volume, natural falling ending. These are the cheapest points in the section and candidates lose them by being tentative on the first item of the test.
Items 3 to 5
One longer clause or a simple two part sentence, eight to eleven words.
Find the single phrase boundary. Deliver two chunks with one small break between them.
Item 6
Two clauses joined by a connector such as "if" or "when".
Identify the connector while listening. It marks the join, so it tells you where the chunks are before you have to speak.
Item 7
The longest, often three parts with a contrast such as "though".
Meaning first. If you understood it as a picture, three chunks come back. If you tried to hold words, the last chunk will be missing.
What to do differently as the block progresses.
Record two: the long one
Repeat this three chunk sentence exactly: "Food and drinks are not permitted near the shelves, though sealed water bottles may be carried through the reading room."
This is an item 7 length sentence. In the first second after the audio ends, do not reach for words. Reach for the picture: a rule about food, an exception about water, and where the exception applies. Then speak. Start the first chunk within about a second and a half and do not stop between chunks to check yourself. If you lose a word in the middle, keep the rhythm going and say the sentence through to the end. A complete sentence with one wrong word scores far better than an accurate half.
0:12
Model answer
Three chunks, two breaks. The model breaks after "shelves" and after "bottles". Those are the two joins in the meaning: rule, exception, scope. Every other word runs on. If your recording has breaks anywhere else, that is a memory join showing through, and it is audible.
"Though" carries weight. The contrast word gets real stress in the model, because it is the hinge the sentence turns on. Candidates almost always underweight connectors, which flattens the logic of the sentence even when every word is present.
The endings survive. Listen to the final consonants on "permitted", "shelves" and "bottles". Under memory strain these are the first things to go, and each one that disappears is a small piece of grammar the listener has to reconstruct.
The pace does not accelerate. A common pattern is starting carefully and then speeding up through the second half as confidence arrives. The model holds one pace throughout. If yours accelerates, it usually means you were reading ahead in memory rather than delivering what you had already secured.
Daily practice that actually moves this block
Ten minutes a day, not an hour on Sunday. Repetition accuracy responds to frequency far more than to session length.
Practise with sentences you have not seen written down. Reading along removes the memory load, which is the whole skill.
Record every attempt and listen back to at least three of them. Judging by feel while speaking is unreliable.
Before repeating, say the number of chunks out loud. Forcing the count builds the habit of hearing structure.
Deliberately practise the last three words of each sentence, since that is where points leak.
Once a week, do all seven items of a practice set in a row with no pausing, to rehearse the real rhythm of the block.
If a sentence beats you twice, listen to it once more for meaning only, then try again. Do not grind it by ear.
Technique check
Three decisions you will actually have to make inside the recording window.
You realise halfway through repeating a long sentence that you have forgotten the final clause. What is the best move?
A well delivered partial sentence is still intelligible and still sits in the middle of the scale. Ending it cleanly protects the delivery dimension, which is scored alongside accuracy, and it keeps you calm for the next item.Silence for the remainder gives a rater nothing and turns a partial response into what sounds like a breakdown. It is the worst of the four options.Restarting costs seconds you may not have and usually produces a second failure at the same point, since the missing clause was never in memory to begin with.Inventing filler introduces words that were not in the source. That is an accuracy error added on top of the omission, and it can push a response from partly correct into hard to understand.
When should you begin speaking after the sentence audio?
A short beat lets the whole sentence settle as a shape without eating meaningfully into the window. It is long enough to prevent the receiving and producing overlap that destroys the back half of long sentences, and short enough that no rater would call it hesitation.Starting before the audio ends is the single most damaging habit in this block. You cannot store the end of a sentence while producing the beginning of it, and the end is where points are lost.Five seconds is most of a short window. On an eight second item that leaves you racing, which produces exactly the rushed, unphrased delivery that costs the top level.Zero gap works on short items and fails on long ones, because it forces you to begin before the shape has formed. Build one habit that works on all seven items rather than two that each work on half.
Which practice method builds this skill fastest?
The scored skill is hold and reproduce. Only a listen once, repeat, then check cycle trains the holding step, and the check afterwards is what turns a failed attempt into information rather than frustration.Reading along removes the memory load entirely, which is the part that is actually being tested. It is useful for pronunciation practice and useless for this task.The sentences on test day will not be sentences you memorised. Memorising a list trains recall of specific content rather than the general ability to hold a new sentence.Isolated sound drills help a narrow pronunciation problem but do nothing for the memory and phrasing issues that cause most of the lost points in this block.
What to expect as you improve
Progress on this block is unusually visible, which makes it good for morale as well as for points. In the first week of daily practice most students find that the short items become automatic and the long ones still collapse. In the second week the collapse point moves later in the sentence, which is the sign that chunking is starting to work. By the third week the long sentences come back complete but the delivery is still slightly flat, and that is the moment to move on to the next lesson, which is about stress, rhythm and intonation specifically.
If you are stuck, the diagnosis is almost always one of two things. Either you are starting too early, in which case the recording will show the second half degrading while the first half is perfect. Or you did not understand the sentence, in which case the recording will show hesitation everywhere rather than at a particular point. The first is a habit and takes a few days to fix. The second is a listening issue, and the fastest treatment is more Listening section practice rather than more repetition practice.
Monthly & Premium
11 more sections in this lesson
You have read the opening. The rest covers the method in full, with worked examples and practice questions that explain why each wrong answer is wrong.
All 45 lessons: Reading, Listening, Writing and Speaking
Real exam audio and record-yourself speaking drills
Seven, and they are items 1 to 7 of the Speaking section. Each carries five points, so the block is worth 35 of the 55 raw points in Speaking, which is more than the four interview answers combined.
No. Each sentence plays once and the recording window opens once. There is no replay control, which is why the technique is built around holding meaning rather than trying to catch every sound.
The window scales with the sentence. In our practice tests it runs from around eight seconds on the shortest early items to about twelve on the longest later ones. The window is generous relative to the sentence, so time pressure is rarely what costs the point.
Say the sentence through anyway with the word you think you heard, and keep the rhythm. One wrong word inside an otherwise fluent, complete repetition sits near the top of the scale. Stopping to correct yourself, or going silent, drops you much further.
No. You are not being asked to imitate the speaker, only to reproduce the sentence with natural stress and rhythm. Your own accent is not an error. Copying the speaker deliberately can even hurt, because attention spent on imitation is attention not spent on holding the sentence.
You cannot, and you should not practise that way either. Writing splits your attention and removes the memory work that the task is measuring. Practise the way you will be tested: listen once, hold it, say it back.