The audio is not too fast, it is connected. The reductions that cost the most marks, US and UK voices, and five clips with full transcripts to train on.
12 min readMonthly & Premium
Quick answer
When TOEFL audio feels too fast, the problem is usually connected speech rather than speed. Words link, unstressed vowels reduce, and sounds shift to match their neighbours. Train on more than one accent, learn the twelve reductions that carry grammatical meaning, and never practise at reduced playback speed.
It is not fast, it is connected
Almost every student who says the listening is too fast is describing something else. Measure the words per minute in a shipped academic talk and you get a rate an average lecturer would consider unhurried. What makes it feel fast is that the words are not separated.
In careful speech, the kind used in a classroom exercise, each word arrives with a clean boundary. In ordinary speech, boundaries dissolve. A final consonant slides onto the next word's opening vowel. Unstressed vowels collapse towards a single neutral sound. Sounds shift to match whatever is next to them. The result is that a sentence you would read without effort becomes a sequence of shapes your ear has never been trained on.
This matters for your score in a very specific way. The words that reduce most are the small grammatical ones: have, had, would, to, for, of, and, can, could, did. Those are exactly the words that carry tense, modality and negation. Miss would have and you lose the fact that something did not happen. Miss the difference between can and can't and you have inverted the meaning of the sentence. Content words like cathedral or urchin are long, stressed and easy to catch; the words that decide the answer are short, unstressed and disappear.
The good news is that connected speech is systematic. It is not sloppiness and it is not random. There are a small number of patterns, they repeat constantly, and once your ear knows them the audio stops sounding fast and starts sounding normal.
The reductions that cost the most marks
Written
What it sounds like
What you lose if you miss it
would have / could have / should have
would've, could've, should've, often shortened further to a single unstressed syllable
That the event did not happen. This is the single most damaging miss in Listening.
did not have to
didn't hafta
That the obligation was absent, not that something was forbidden.
can't
the vowel shortens and the final t often disappears entirely, leaving stress as the only clue
The negative. Stress is the real signal here: cannot is stressed, can is not.
going to
gonna
Nothing, if you know it. A surprising number of learners still parse this as two separate words and stall.
want to / have got to
wanna, gotta
Speed only, but the stall while you decode costs you the next clause.
and
n, often attached to the previous word
List boundaries, which matters when an announcement lists three required documents.
of / for / to
a neutral unstressed vowel, barely audible
Relationships between nouns, which is where inference questions live.
half eight, quarter past
British time expressions spoken quickly
The exact time in an announcement, which is the most commonly tested detail there is.
Each of these hides a grammatical distinction. Learn to hear the shape rather than the spelling.
Clip 1: a modal perfect at speed
What does the speaker mean?
I would have gone means the going did not happen. The whole meaning of the sentence sits in a contracted, unstressed form that lasts a fraction of a second, and the rest of the clip is the reason it did not happen.This is what you hear if the modal perfect goes past you and you keep only gone to the seminar. Note that every content word in the sentence points this way; only the grammar says otherwise.Nothing about a future seminar is said. Learners often add a hopeful future when they sense a negative but cannot locate it.The opposite of couldn't get out of my shift in time. Missing the negative in couldn't reverses the second clause too.
Show transcript
I'd have gone to the seminar, but I couldn't get out of my shift in time.
Clip 2: absence of obligation
What is the speaker telling the listener?
Did not have to marks an absent obligation, and the past tense plus you know tells you the listener already did the unnecessary thing. Two small unstressed words carry the entire situation.This is the classic confusion between did not have to and must not. One says there was no need, the other says it was forbidden, and the audio gives you no more than a syllable to tell them apart.No instruction is given. The cart is mentioned as the reason the effort was unnecessary, not as a next step.A rule about who may use the cart is invented here. The speaker only says the laptops exist.
Show transcript
You didn't have to bring your own laptop, you know. There's a whole cart of them in the room.
Clip 3: a British announcement with fast numbers
When does the library reopen?
Half eight in British usage means half past eight, not half before eight. It is spoken in two syllables and goes past very quickly, which is why time expressions belong on your paper as numerals the instant you hear them.This is the reading a speaker of German or Dutch would expect, where the equivalent phrase means half before the hour. In British English it means half after.Eight o'clock drops the half entirely, which is what happens when the phrase is heard as a single blurred unit.Wednesday is in the clip, attached to the coursework deadline rather than the reopening. Two dates in one announcement is the standard trap.
What are students told to do about the coursework deadline?
The speaker says you do not need to email anyone about it and that the deadline has been moved for everybody. The whole instruction is a negative one, and the negative is unstressed.This is exactly what the announcement rules out. Options that describe the action a student would normally take are the most tempting kind of wrong answer in announcements.Early submission is never mentioned. The deadline moves later, which removes any reason to hurry.No form of any kind appears in the clip. This option borrows the library reopening and attaches an invented errand to it.
Show transcript
Right, a few notices before you go. The library will be shut for the whole of the bank holiday weekend, reopening at half eight on Tuesday. Coursework that was due on the Monday now goes in on the Wednesday, and you don't need to email anyone about it, the deadline's been moved for everybody. If you've booked a study room over the weekend, that booking has been cancelled automatically, and you'll see the deposit refunded within a fortnight.
A shadowing routine that actually changes your ear
1
Choose thirty seconds, not five minutes
Take a short stretch of audio with a transcript. Thirty seconds is enough. The purpose is not exposure, which you get from ordinary listening; it is precision, and precision needs repetition on a small target.
2
Listen three times without the transcript and write what you hear
Write it as words, gaps included. Leave a blank where you heard a shape but not a word. Those blanks are your actual syllabus. Most of them will be the small grammatical words in the table above, which is the point.
3
Open the transcript and mark every gap
Now you can see what the shape was. Circle the joins: where a consonant ran into a vowel, where a vowel reduced, where a t vanished. After a fortnight of this you will start recognising the same half dozen joins over and over, because there are only a few and they recur constantly.
4
Speak along with the audio, twice
Say it at the same time as the speaker, matching the rhythm rather than the individual sounds. This is the part people skip and it is the part that works. Producing the reductions yourself is what makes your ear stop expecting the written form. You do not need a good accent for this; you need the same timing.
5
Come back to the same clip two days later
Listen once, cold. What used to be a blur will be words. That moment is the evidence that the method works, and it is worth engineering deliberately, because listening practice usually gives you no visible progress at all.
Clip 4: a fast conversation with contractions throughout
Why had Speaker B decided not to attend?
His first turn gives the reason directly, and the key part is I have not started it. Note that I was going to, with the verb left off, is a complete answer in spoken English; the listener supplies go from context.The big lecture hall appears later, in his mistaken belief about which sessions get recorded. Reusing a real phrase from a different part of the clip is the standard build for this kind of distractor.He clearly knows about it and had intended to go. I was going to establishes prior intent in four unstressed syllables.The preference for the recording arrives only after he learns it exists. Order matters: it is the solution, not the original reason.
What does Speaker A warn Speaker B about?
The warning is the second-to-last line, and like most warnings in a conversation it is compressed: do not leave it more than a week, then the reason. Late-clip information is where next-step and warning questions consistently come from.That was Speaker B's incorrect assumption, and Speaker A corrects it. Attributing a belief to the person who disproved it is an attribution error.Speaker A says the opposite: it goes up the same evening, around nine.Attendance rules are never discussed. This is a plausible university policy imported from outside the clip.
Show transcript
Speaker A: Are you going to the info session tomorrow?
Speaker B: I was going to, but I've got a lab report due at midnight and I haven't started it.
Speaker A: They're recording it, though. You could just watch it afterwards.
Speaker B: Could I? I thought they only recorded the ones in the big lecture hall.
Speaker A: They record all of them now. It goes up the same evening, usually around nine.
Speaker B: That's perfect, actually. I'll finish the report first and watch it while I eat.
Speaker A: Just don't leave it more than a week. They take them down after seven days.
Speaker B: Good to know. I'll do it tomorrow night.
Clip 5: a linguistics talk about the thing you are practising
According to the speaker, what does reduction do?
Stated in the middle of the numbered list. When a lecturer enumerates three processes and defines each, the fact question almost always asks you to attach one definition to the right label.Consonant deletion is not one of the three processes named, though it happens in real speech. Options built from true general knowledge that the clip did not state are common on lists.That is assimilation, the third process. Swapping items within a list is the standard difficulty in this question type.That is linking, the first process. Again a real definition attached to the wrong label.
What can be inferred about asking a speaker to talk more clearly?
The speaker says careful speech is a different register from the one you will actually meet, and that you cannot close the gap that way. The inference is the combination of those two statements, which is why the answer feels obvious once assembled and is not quoted anywhere.Politeness is never raised. The argument is about which variety of speech you are training on, not about social cost.No distinction between lectures and conversations appears in this part of the talk. Splitting a general claim into contexts the speaker did not mention is a frequent inference trap.The clip argues the opposite. It says this route does not work, so calling it the fastest improvement reverses the point.
What will the speaker discuss next?
The final sentence announces it outright. Every academic talk in this course closes with a forward pointer for the same reason the test does: real lecturers signal what is coming.Other languages are mentioned in passing, to say these processes occur in all of them. That is a supporting point, not an announced next topic.Teaching adults is close to the practical consequence he drew, but he draws that consequence and then moves on rather than promising to develop it.Careful speech is defined at the start as a contrast. It is finished business by the end of the clip.
Show transcript
When people say a language sounds fast, what they usually mean is that it's connected. In careful speech we produce words with clear boundaries. In ordinary speech we don't, and three processes do most of the work. The first is linking, where a final consonant attaches to the next word's opening vowel, so pick it up comes out as a single unit. The second is reduction: unstressed vowels collapse towards a neutral sound, which is why to, for, of and and almost never sound the way they look on the page. And the third is assimilation, where a sound shifts to match the one beside it, so that ten pounds ends up much closer to tem pounds. Notice that none of this is carelessness. These processes are systematic, they occur in every language, and speakers apply them without noticing. That has a practical consequence for anyone learning to listen: you can't close this gap by asking people to speak more clearly, because careful speech is a different register from the one you'll actually meet. Next time, we'll look at how children acquire these patterns years before they can read.
Your accent and speed programme
Practise on both US and UK voices every week. Our lesson audio deliberately alternates between them for this reason.
Never reduce playback speed. Use the transcript to close the gap instead.
Do one thirty-second shadowing block four times a week. Small target, high repetition.
Keep a running list of the reductions that caught you. It will be short and it will repeat.
Practise the modal perfects specifically: would have, could have, should have. They carry more meaning per syllable than anything else in the section.
Learn the British time expressions, half eight and quarter past, because announcements test times more than any other detail.
Listen to something unscripted, not just test material. Unscripted speech has the false starts and repairs that scripted audio smooths away.
What four weeks of this actually looks like
Week one, you will do the dictation step and find that a third of your blanks are the same handful of function words. That is discouraging for about a day and then it becomes the most useful information you have, because a short list of specific failures is a plan and a general sense of weak listening is not.
Week two, the shadowing starts to feel less absurd and your timing improves before your accuracy does. Do not chase accuracy here. Matching the rhythm is what retrains expectation; matching the individual sounds is a pronunciation goal and belongs to the Speaking course.
Week three, you will notice something odd in ordinary listening: you catch a reduced would have in a film without trying. That is the transfer you are working for. It shows up outside practice before it shows up inside it, because outside practice you are relaxed.
Week four, return to a practice module you took at the start and count the items you now get for a different reason than you did before. Not more items necessarily, but a different set: the ones that turned on a negative, a modal or a time expression.
One realistic caution. This work improves comprehension of fast speech, which is a large slice of Listening but not all of it. If you are also losing items to distractors, to attribution, or to fatigue in the last clip of a module, those are separate problems with separate lessons in this course. Diagnose before you train, or you will spend four weeks fixing something that was not the thing costing you the band.
Monthly & Premium
9 more sections in this lesson
You have read the opening. The rest covers the method in full, with worked examples and practice questions that explain why each wrong answer is wrong.
All 45 lessons: Reading, Listening, Writing and Speaking
Real exam audio and record-yourself speaking drills
Our practice audio uses both US and UK voices, and the safe assumption is that you should be comfortable with more than one variety of English rather than only the one you have studied most.
No. Measured in words per minute it sits at ordinary lecture and conversation pace. What makes it feel fast is connected speech: linking, reduction and assimilation removing the word boundaries you rely on when reading.
No. Slowing it removes exactly the features you need to learn and trains your ear on a register nobody uses. Listen at full speed, use the transcript to see what you missed, then listen again.
The modal perfects, would have, could have and should have, because they tell you an event did not happen. After those, negatives such as could not and did not have to, and the small function words of, for, to and and.
Most students report a clear change after three to four weeks of short daily dictation and shadowing work. It arrives in ordinary listening before it arrives in test conditions, which is a good sign rather than a strange one.