Word-perfect repetitions still score four out of five when delivery is flat. The sentence stress, weak forms, linking and intonation that reach the top.
11 min readMonthly & Premium
Quick answer
The top level of Listen and Repeat asks for perfect repetition with natural intonation and rhythm, so accurate words alone reach four out of five, not five. English rhythm comes from stressing content words, compressing function words, linking across word boundaries, and letting pitch fall at the end.
The gap between correct and natural
You reproduced all fourteen words. Every sound was clear. The score came back as four out of five, and the feedback said something about rhythm. This is the most common frustration in the whole Speaking section, and it is not arbitrary. Word accuracy gets you to the second level of the scale. The top level explicitly requires natural intonation and rhythm on top of accuracy, and there is a reason the scale is built that way.
Speech carries meaning in two channels at once. The words say what happened. The music says which part matters, where the sentence divides, whether it is finished, and how the speaker feels about it. A listener uses both channels without noticing, and when the second channel is missing they have to do the work themselves. That extra work is exactly what the difference between a five and a four measures.
The good news is that English rhythm is rule governed and small. There are about four things to learn, they interlock, and once they click they apply to every sentence you will ever say, in the interview block as much as in the repetition block. This lesson covers them in the order that produces the fastest audible change.
Rule one: English gives time to meaning, not to words
In some languages every syllable takes roughly the same amount of time. English is not one of them. English gives real time to the words that carry meaning and squeezes everything else into the gaps between them. The words that get time are nouns, main verbs, adjectives, adverbs and question words. The words that get squeezed are articles, prepositions, auxiliary verbs, pronouns and conjunctions.
Take "The study carrels are located on the quiet upper floor." Five words carry the meaning: study, carrels, located, quiet, upper, floor. Everything else is scaffolding. A natural delivery lands hard on those content words and races through "the", "are", "on the" so fast that they almost disappear. A flat delivery gives all eleven words the same weight, and the result sounds like a list rather than a sentence.
This single change does more for how natural you sound than any amount of individual sound work. It is also counterintuitive for many learners, because careful pronunciation of every word feels like good pronunciation. In English it is the opposite: compressing the small words is what makes the big ones audible.
Content words versus function words
In the natural version, which words are almost inaudible?
Articles, auxiliaries and prepositions are the compressed class in English. They keep their grammatical job but lose almost all of their time and their vowel quality, which is what creates the gaps that the content words fill."Study" and "carrels" are the nouns naming the thing under discussion, so they take some of the heaviest stress in the sentence."Quiet" and "upper" are adjectives that distinguish this floor from other floors. They are exactly the kind of information English protects with stress."Located" is the main verb and "floor" is the final content word, which typically carries the sentence's last strong beat before the pitch falls.
Why is the natural version physically shorter than the word by word version?
Nothing was cut and nothing was rushed. The unstressed words simply took the reduced form that English uses by default, which takes far less time than the careful citation form of each word.The content words in the natural version are not faster. Several of them are actually longer than in the word by word version, because they are carrying the beats.Both versions contain the identical ten words. That is what makes the comparison useful: the only variable is timing.The vowels in the unstressed words get shorter, not longer, and the pauses between words disappear rather than being replaced by anything.
Show transcript
Listen to what happens when every word gets the same weight. The. Study. Carrels. Are. Located. On. The. Quiet. Upper. Floor. Now listen to the same sentence with the meaning words carrying the time. The study carrels are located on the quiet upper floor. The second version has exactly the same words. It is shorter, and it is much easier to follow.
The words English squeezes
Word
Careful form
What it becomes in running speech
to
too
a very short "tuh", as in "going tuh the desk"
for
for, with a full vowel
a quick "fer", as in "fer up to three weeks"
and
and
often just "n", as in "food n drinks"
are
are
reduced to a small "er" or attached to the word before it
can
can
a short "kn", clearly different from the stressed "can't"
at the
at the
run together into one compressed unit, "atthuh"
of
of
a light "uv", almost lost between two content words
Function words and what they usually sound like inside a normal sentence.
Rule two: words join up
English does not leave gaps between words. Inside a phrase, the end of one word attaches to the start of the next, and the joins follow patterns. A final consonant moves onto a following vowel, so "keep it" becomes "kee-pit" and "check out" becomes "che-kout". Two identical consonants merge into one longer one, so "want to" becomes "wanna" in casual speech and at minimum loses one of the two t sounds. A word ending in a vowel and one starting with a vowel get a small glide between them.
Learners who articulate every word separately sound careful and score as unnatural. The joins are not sloppiness, they are the structure. When you hear a native sentence and think it went too fast, what you are usually hearing is not speed but linking: fewer boundaries than you expected, so your ear cannot find the word edges it was listening for.
This matters twice over in this section. Producing the joins is what makes your repetition sound natural. Expecting the joins is what lets you hear the sentence correctly in the first place, which is why linking practice quietly improves your Listening score as well.
Building a natural repetition, layer by layer
1
Layer one: find the thought groups
Divide the sentence where the meaning divides, usually two or three groups. Each group is delivered as one connected run with no internal stops. The break between groups is tiny, often less than a quarter of a second, but it is the thing that makes a long sentence sound organised rather than endless.
2
Layer two: pick one strongest word per group
Every thought group has a peak, and it is normally the last content word in the group. In "Food and drinks are not permitted near the shelves", the peak is "shelves". Hitting one clear peak per group is what gives English its wave shape, and it is far more effective than trying to stress everything important.
3
Layer three: compress everything that is not carrying meaning
Actively shorten the articles, prepositions and auxiliaries. This feels wrong at first, like you are being careless. Record it and listen: it sounds like fluency, not carelessness. This is the layer that most obviously separates a four from a five.
4
Layer four: let the pitch fall at the end
A statement ends with the voice going down. If your pitch stays level or rises at the end of a statement, the listener hears an unfinished sentence or an unintended question, and the response reads as incomplete even though every word was there.
5
Layer five: check the multi-syllable words
Words like "permitted", "available", "convenient" and "photography" carry stress on one specific syllable, and putting it elsewhere can genuinely obscure the word. This is the only layer that is memorisation rather than pattern, so learn the ones that appear in campus contexts and let the rest come with exposure.
Record one: three thought groups, three peaks
Repeat this sentence with deliberate attention to rhythm: "The self-checkout machine accepts student cards, but guest passes must be validated at the desk."
Before you speak, mark the groups in your head: machine accepts cards, then the contrast, then where validation happens. In the first second after the audio ends, begin the first group as one connected run. Do not pause inside a group even if a word feels shaky. Aim for exactly two small breaks in the whole sentence, both at the group joins, and let your pitch drop clearly on the final word.
0:12
Model answer
Three peaks, not eleven. The model puts real weight on "cards", "passes" and "desk", with secondary weight on "self-checkout" and "validated". Everything else is scaffolding. Count the strong beats in your own recording: if there are more than five, you are giving time to words that do not need it.
"But" marks the turn. The contrast word gets a small lift and a slight break before it. That single move tells the listener the sentence is about to reverse, and it is the clearest signal of comprehension a rater can hear.
The compressions. Listen to "must be" and "at the" in the model. They are fast, low and short. If those four words take as long in your version as "validated" does, the sentence will sound like a list.
The fall. "Desk" drops. If your version holds the pitch level on the last word, the sentence sounds like the first half of something longer, and that costs a point that has nothing to do with accuracy.
Intonation changes the meaning, not just the mood
A candidate repeats a statement but ends with rising pitch. What does a listener hear?
Rising pitch at the end of a statement is the English signal for either a question or more to come. The listener waits for the rest, then has to reinterpret, and that moment of extra work is what the delivery part of the scale is measuring.This is not an accent feature. Speakers of every English accent fall at the end of a plain statement, so the rise reads as a delivery choice rather than a regional one.A final rise can soften a request in some contexts, but on a flat declarative repetition it signals incompleteness, not politeness.The words are identical and the meaning conveyed is not. That is precisely the point: intonation carries information the words do not.
Why does the list example fall on the final item?
The rise, rise, fall pattern is how English marks the boundary of a list. The fall tells the listener nothing further is coming, which is why a list delivered with all rises sounds like it was interrupted.Importance is not what the pattern encodes. Reorder the three items and the fall moves to whichever one is last, regardless of which matters most.Word length has nothing to do with it. "Maps" is the shortest of the three and still takes the fall because of its position.Volume is not the mechanism. The final item is not quieter, its pitch contour goes down, which is a different thing and the one that carries the meaning.
Show transcript
The same words can do different jobs. Listen. The library is open until nine. That is a statement, and the voice falls at the end. The library is open until nine? That is a question, and the voice rises. Now a list. The desk has forms, passes, and maps. The voice rises on forms and passes, and falls on maps, which is how you know the list has finished.
Record two: a list and a close
Repeat this sentence, paying attention to the list contour: "You will need your student card, a completed form, and proof of enrolment before the pass can be issued."
This sentence has a three item list inside it, then a closing clause. In the first second, fix the three items in mind as a set rather than as separate words. Then deliver them with the voice lifting on the first two and settling on the third, and take the final clause down to a clear finish. Do not stop between the list and the closing clause for longer than a beat.
0:12
Model answer
The list is audibly a list. The model lifts on "card" and on "form", then settles on "enrolment". A rater can hear the structure of the sentence without parsing the words, and that is the strongest possible evidence that you understood what you repeated.
The three items are parallel in rhythm. Each one gets roughly the same shape and the same time. If one item comes out much slower than the others, that is retrieval time showing, and it breaks the pattern the listener is tracking.
The final clause is subordinate and sounds it. "Before the pass can be issued" is delivered lower and faster than the list, because it is background information. Giving it the same weight as the list would flatten the sentence into an undifferentiated string.
One clear ending. The pitch falls decisively on "issued". Listen for whether yours does, and specifically whether the final consonant is fully formed rather than trailing into breath.
A rhythm audit for any recording
I can name the thought groups in the sentence and my breaks fall only at those joins.
Each thought group has one clearly strongest word, not three competing ones.
My articles, prepositions and auxiliaries are noticeably shorter than my nouns and verbs.
Words inside a group are joined, not separated by small gaps.
The pitch falls at the end of the sentence and the final consonant is fully pronounced.
Any list inside the sentence rises on the early items and falls on the last.
Contrast words such as "but", "though" and "however" get a small lift.
My pace is even throughout rather than accelerating as I gain confidence.
Delivery decisions
Four questions on the choices that separate the top two levels of this block.
Which sentence would most likely be marked down for rhythm despite perfect word accuracy?
Equal weight on every word is the definition of the flat delivery that the scale distinguishes from natural rhythm. No error is present, and no English rhythm is present either, which is exactly the four out of five profile.Stressing content words and compressing function words is the core of natural English rhythm, so this is the delivery the top level is describing.Small breaks at meaning joins is correct phrasing. It is what makes a long sentence parseable and it is rewarded, not penalised.A falling ending on a statement is the expected pattern and signals a complete response.
You are unsure which syllable of "available" carries the stress. What should you do inside the recording window?
You just heard the word spoken correctly moments ago. The audio is the answer key, and copying what you heard is both the fastest route and the one the task is designed around.Substituting a synonym introduces a word accuracy error into a task where reproducing the exact sentence is half the score.Even stress across four syllables is the flat delivery pattern, and on a long word it can genuinely obscure which word you meant.A pause to think costs rhythm and inserts a break at a point where there is no meaning boundary, which is one of the clearest markers of strained delivery.
What is the practical benefit of practising linking between words?
Linking is a two way skill. Producing the joins makes your delivery sound native, and expecting the joins is what lets you correctly parse a sentence that seemed too fast, which helps in both Speaking and Listening.Speed is not the goal and the window is not the constraint on these items. Linking changes how words connect, not how quickly you deliver content.Linking moves final consonants onto the next word, it does not delete them. Dropping final consonants is a separate habit and it costs points.Linking applies to all speech. It matters in every one of the eleven items, and the repetition block is where it is most directly scored.
A two week rhythm plan
Days one to three, work on compression only. Take five sentences a day, mark the function words, and practise making them short. Ignore everything else. Record and listen specifically for whether "the", "to", "and" and "of" are noticeably shorter than the words around them.
Days four to seven, add thought groups. Same five sentences a day, but now mark the groups first and deliver each group as one unbroken run. The test is whether your breaks land only at the joins. Most people find one habitual break that does not belong to any meaning boundary, usually just before a long word, and killing that single break is worth a visible amount.
Days eight to eleven, add the endings. Falling pitch on statements, full final consonants, no trailing off. This is the fastest layer to fix and the one candidates most often neglect, because the mistake happens after the interesting part of the sentence and attention has already moved on.
Days twelve to fourteen, put it together at full speed on complete practice sets, seven items in a row with no stopping. By this point you should not be thinking about any of the four layers consciously. If you still are, do another week of layer three rather than pushing on. Compression is the one that needs to become automatic before the rest can.
Monthly & Premium
12 more sections in this lesson
You have read the opening. The rest covers the method in full, with worked examples and practice questions that explain why each wrong answer is wrong.
All 45 lessons: Reading, Listening, Writing and Speaking
Real exam audio and record-yourself speaking drills
Because the top level of the scale asks for perfect repetition with natural intonation and rhythm, not just correct words. A flat delivery with equal weight on every word is accurate and unnatural at the same time, and that combination is what the second level of the scale describes.
No. Nothing in the scale asks for a particular accent, and imitating one is a distraction that costs attention you need for holding the sentence. What is scored is whether your stress, rhythm and intonation follow English patterns, and those are the same across accents.
Content words carry the meaning: nouns, main verbs, adjectives, adverbs and question words. Function words hold the grammar together: articles, prepositions, auxiliaries, pronouns and conjunctions. English gives real time to the first group and compresses the second, and that contrast is what creates the rhythm.
On the TOEFL it works against you. Full careful pronunciation of function words produces the even, listy delivery that the scale separates from natural speech. Compressing them feels careless and sounds fluent, which is the opposite of what most learners expect.
Hum it. Listen to the model, then hum the tune with no words, then hum your own recording. Any mismatch in shape is immediately obvious to your ear even when you cannot describe it in technical terms, and it isolates the music from the words you already know you got right.
No. Delivery is part of what a listener judges on every interview answer as well, and monotone is one of the quietest ways to lose points there because it produces no identifiable error. Rhythm work you do for items 1 to 7 pays out again on items 8 to 11.