How Writing Is Scored: The Rubric in Plain English
What separates a 3 from a 4 and a 4 from a 5 on TOEFL Writing. The 0 to 5 band descriptors for the email and discussion tasks, decoded with real examples.
11 min readMonthly & Premium
Quick answer
Each TOEFL Writing task is scored 0 to 5. Task completion decides the band first: if a required point is missing or vague you are capped at 3 or below, however good your English is. Language quality then separates a 3 from a 4 and a 4 from a 5.
Three scores, one band
Writing produces three numbers before it produces a band. Your ten Build a Sentence items give a total out of 10, which converts to a score out of 5. Your email is scored 0 to 5. Your discussion post is scored 0 to 5. Those three are averaged, and the average is what drives your Writing band, which runs from 1.0 to 6.0 in half-point steps and is aligned to CEFR levels.
Two consequences follow, and both matter more than any writing tip in this course.
First, Build a Sentence is not a warm-up. It carries the same weight as the entire email. A student who takes 6 out of 10 there has already given away as much as a student who wrote a band 2 email.
Second, averaging punishes gaps. Two strong components and one weak one produce a mediocre band. This is why the pacing rule in this course is always the same: get something complete on the screen for every task, even if the last two sentences are plainer than you would like.
The email rubric, decoded
Band
The descriptor says
In practice this means
5
Fully addresses all three prompt points, clear purpose, appropriate greeting and closing, consistently polite tone, sentence variety, accurate grammar, precise vocabulary.
All three points genuinely developed, not just named. A greeting and a signed closing that fit the reader, formal for an office or a professor and first-name for a peer. A mix of short and longer sentences. Errors are rare enough that you have to look for them.
4
Addresses all three points adequately, generally appropriate tone, minor grammar or word-choice issues do not impede understanding, logical organisation, greeting and closing present.
Everything is covered and the reader is never confused, but a few articles are missing, a preposition is off, or one sentence is clumsy. This is the most common good score.
3
Addresses most points but one may be vague or underdeveloped. Tone mostly suitable. Noticeable grammar or vocabulary problems, message still understandable.
You mentioned the third point in six words at the end, or your request was so general the reader would not know what to do. Fixing the content, not the grammar, is what moves this to a 4.
2
Partially addresses the prompt. Notable problems with completeness, register or grammar. Missing greeting or closing, or one point omitted.
A whole required point is absent, or you opened with Hi and stopped writing without a sign-off. Both of those are avoidable in fifteen seconds of checking.
1
Minimally addresses the situation. Serious language errors make portions hard to understand, or two or more required points are missing.
Usually a time problem, not a language problem. The response was cut off before two of the three points were reached.
0
Off-topic, blank, copied verbatim from the prompt, or written in a language other than English.
Copying the situation paragraph as your opening is the realistic risk here. Paraphrase the situation, never lift it.
Left column is the official descriptor in condensed form. Right column is what it means when you are looking at your own draft.
The academic discussion rubric, decoded
Band
The descriptor says
In practice this means
5
A clear, well-elaborated contribution with a strong opinion that engages meaningfully with the discussion. Well organised and coherent, varied sentence structure, accurate grammar, precise vocabulary.
Your position is unmistakable, you answer something a named classmate actually said, and you develop one reason with a specific example rather than listing three shallow ones.
4
Relevant contribution with an opinion supported by reasons or examples. Adequate development and organisation. Occasional minor errors do not obscure meaning.
Good post, clear opinion, real support, but the classmates are ignored or acknowledged only in a generic way, or the example stays abstract.
3
Mostly on topic but may lack depth or specific examples. Some organisational issues, noticeable errors, meaning generally clear.
You said what you think and gave a reason, but the reason is the kind of thing anyone could write without reading the prompt. This is the ceiling for generic opinion writing.
2
Limited relevance or development, weak organisation, frequent errors that sometimes obscure meaning.
Often a post that drifts off the actual question, or answers a simpler version of it.
1
Largely irrelevant, undeveloped or incoherent, with severe and persistent language errors.
Rare unless the response was abandoned early.
0
Blank, off-topic, not in English, or copied from the prompt.
Also applies when the post is a single opening sentence followed by nothing, because no dimension can be assessed.
Same structure, different priorities. Engagement replaces required points.
The band discipline rule that surprises people
A response full of errors is not automatically a 1 or a 2. If the student covered every required point and the message remains understandable overall, the descriptors put the floor at 3, no matter how dense the spelling and grammar problems are. Errors that genuinely obscure meaning are what push a complete response down to a 2.
This cuts both ways, and the second direction is the one that costs students marks. Immaculate English does not lift an incomplete response. A flawless two-paragraph email that never asks the question the prompt told you to ask is a 3, and it will stay a 3 no matter how many times you polish the sentences.
If you take one habit from this lesson, take this one: before you check your draft for grammar, count the required points and put a mental tick against each one. That check takes ten seconds and it protects the criterion that decides your band first.
How to score your own practice response honestly
1
Step 1. Count the required points
For the email, the prompt gives you three explicit instructions. Find the sentences in your draft that do each one. If you cannot point at a specific sentence, that point is not covered, and no amount of good writing elsewhere fixes it. For the discussion, the equivalent check is: can I point at my position, at my engagement with a named classmate, and at my supporting example?
2
Step 2. Check the frame
Greeting present and appropriately formal. Closing present. Your name signed. For the discussion, a closing sentence rather than a mid-thought stop. These are worth a band on their own and take no skill, only memory.
3
Step 3. Read for meaning-blocking errors only
Not every error, only the ones that make a reader stop and reread. A missing article rarely does. A tense that contradicts the situation, a pronoun with no clear referent, or a run-on that fuses two ideas often does. Those are the errors that separate a 4 from a 3.
4
Step 4. Read for register
Register is judged against the reader the prompt gave you, so check the address line first. Most of these emails go to a professor, an office or a manager, and there contractions, casual openers and phrases like thanks a lot belong in a message to a friend rather than here. A minority go to a classmate, a roommate or a tenant, and there the warm version is the correct one and contractions are not marked down. A single slip in either direction will not sink you, but a tone that stays wrong for the reader is explicitly named in the band descriptors as a register problem.
5
Step 5. Only now, look at range
Sentence variety and precise vocabulary are what separate a 4 from a 5, and they are the last thing to work on because they are worth nothing if steps 1 to 4 have failed.
How Build a Sentence is scored
Build a Sentence is scored mechanically. Each item is one point, and you either produce an accepted ordering of the chunks or you do not. There is no partial credit for getting five of six blanks right.
One important detail: a good number of items have more than one correct answer. Word order in English is often flexible, and an item whose chunks include both a present and a past verb form, or a mobile adverb like already, first, just or mainly, can legitimately be arranged in two different grammatical ways. Both earn the point.
On our practice tests those alternatives are stored on the item itself and the scorer credits any of them, so if you build a genuinely correct English sentence that is not the one we happened to list first, you still get the mark. If you ever meet an item where you are confident your ordering is standard English and it was marked wrong, that is worth reporting, because it is a scoring bug rather than a grammar disagreement.
Fast self-check before you call a practice response finished
Every required point has a sentence I can point at, not just a mention.
Greeting and closing are present, and my name is on the email.
If the email goes to a professor, an office or a manager, there are no contractions anywhere in it. If it goes to a peer, contractions are fine and I have not stiffened the tone to avoid them.
The discussion post names at least one classmate and answers something they actually argued.
My example is specific enough that it could not have been written before I read the prompt.
The final sentence is a complete sentence, not a cutoff.
I have at least one sentence longer than fifteen words and at least one shorter than eight.
Where did the marks actually go?
Three real score profiles. In each case decide where the recoverable marks are, not where the lowest number is. Those are usually different places, and mistaking one for the other is why students spend months on the wrong section.
Student A, Writing section
Build a Sentence: 6 out of 10 Email: 4 out of 5 Academic discussion: 4 out of 5
Where should this student spend their preparation time?
Build a Sentence is converted to a score out of 5 and averaged with the other two, so 6 out of 10 enters the average as 3. It is both the lowest component and the one that responds fastest to practice, because it tests a closed set of recurring grammar patterns rather than a general writing ability. Four missed items is the largest single block of recoverable marks in this profile.The discussion does not carry more weight than the other two. Each of the three components contributes equally to the average, which is the fact students most often get wrong about this section.There is a mark available, but the last mark on a task already scoring 4 is the most expensive one on the page. The gap between 6 and 10 on Build a Sentence is cheaper.They are not close together once Build a Sentence is converted. 6 out of 10 becomes 3 out of 5, which puts it a full point below the other two.
Student B, Writing section
Build a Sentence: 10 out of 10 Email: 4 out of 5 Academic discussion: 0 out of 5, left blank when time ran out
What is the most important thing this student should change?
A blank task scores zero and still counts in the average, so it drags the whole section down further than any weak answer could. This student has demonstrated strong language on the other two components, which means the zero is a clock problem, not an ability problem. A rushed, short discussion post scores far above nothing, so the rule is to always produce a finished short answer rather than an unfinished good one.A perfect Build a Sentence score is direct evidence against a grammar problem. The zero came from not writing, not from writing badly.Something structural is very wrong: an entire task went unattempted. The language being fine is exactly what makes that so costly.Vocabulary on the email is worth at most one mark here, and it is unreachable while a whole task is scoring zero.
Student C, Writing section
Build a Sentence: 9 out of 10 Email: 5 out of 5 Academic discussion: 3 out of 5 Feedback on the discussion notes that the two classmates named in the prompt were never mentioned.
What does this profile tell you?
A 5 on the email and 9 on Build a Sentence rule out a language problem. The discussion lost marks because a required move was skipped: the prompt put two named classmates on screen and the post did not engage with either. This is the most common single reason a strong writer stalls at a 3 on this task, and it is worth naming clearly because it is fixable in one sentence rather than in months of language work.Fluency is demonstrably fine. A student who scores 5 on the email is not short of written fluency.Length is not the issue and adding words without engagement would not move the score. The missing element is engagement, not volume.Both tasks are scored out of 5 against their own criteria. The difference here is that one task had a requirement that was not met.
Score these yourself
Five questions. Each one is a judgement a rater makes routinely, so getting them right means you can grade your own practice.
A student writes an email that covers all three required points clearly, but it contains roughly one grammar or spelling error per sentence. Meaning is never lost. What is the lowest defensible band?
The descriptors are task-completion aware. A complete, understandable response sits at 3 or above however dense the errors are. Errors have to obscure meaning before they drag it lower.Band 1 is for responses missing two or more required points or so damaged that parts cannot be understood. Neither is true here.Band 2 requires a completeness, register or framing failure, or errors that begin to obscure meaning. This response has none.Band 4 requires errors that are minor. One error per sentence is more than minor even if meaning survives.
A discussion post takes a clear side, gives two reasons, and is written in near-perfect English. It never mentions either classmate. Where does it realistically sit?
The band 4 descriptor is a relevant contribution with an opinion supported by reasons. Band 5 asks for engagement with the discussion, which is exactly what is missing.Accurate language is necessary for a 5 but not sufficient. The 5 descriptor specifically requires meaningful engagement with the discussion.Band 2 means limited relevance or development with frequent errors. A clear, well-supported, accurate post is nowhere near that.Band 0 is blank, off topic, not in English, or copied. Answering the professor question on topic is none of those.
Which of these changes moves an email from a 3 to a 4 fastest?
The gap between 3 and 4 is defined as one point being vague or underdeveloped versus all points addressed adequately. Fixing the vague point is the whole difference.Vocabulary upgrades affect the 4 to 5 boundary, and forced synonyms often introduce word-choice errors that cost more than they gain.Developing the paragraph that is already working does nothing for the criterion that is capping you, and it eats the time you needed for the weak point.Longer sentences increase your error rate under time pressure. Sentence variety helps, uniform length does not, in either direction.
You arrange a Build a Sentence item into a sentence that is standard English but is not the phrasing you expected the test to want. What normally happens?
Word order in English is often flexible, particularly with mobile adverbs and with verb forms where two tenses are both defensible. Items like that carry a list of accepted orders and any of them earns the point.Scoring against a single fixed key is exactly the bug that was fixed. Items with legitimate alternatives now credit them.Build a Sentence has no partial credit. Each item is one point, all or nothing.Items are not discarded. They are scored against every ordering that is genuinely grammatical for that item.
A student scores 10 out of 10 on Build a Sentence, 4 on the email, and then runs out of time on the discussion and submits two sentences that score 1. Why is the band lower than they expect?
The Build a Sentence total converts to a score out of 5 and is averaged with the email and discussion scores. An average of 5, 4 and 1 is much closer to the middle of the scale than the strong components suggest.Build a Sentence is a full third of the averaged score, not a tiebreaker. That is why it is worth training.There is no double penalty. The single low score is enough on its own because of the averaging.The email and the discussion are each worth 5 raw points. Neither outweighs the other.
What to do with this
Grade three of your own past practice responses using the five-step check above before you write anything new. Most students discover the same thing: their band was decided by completion and framing, not by the grammar they have been worrying about. That discovery is worth more than another practice test.
Then work on the criteria in the order the rubric weighs them. Completion first, framing second, meaning-blocking errors third, register fourth, range last. Working in that order means every hour you spend is spent on the thing currently capping you.
Monthly & Premium
11 more sections in this lesson
You have read the opening. The rest covers the method in full, with worked examples and practice questions that explain why each wrong answer is wrong.
All 45 lessons: Reading, Listening, Writing and Speaking
Real exam audio and record-yourself speaking drills
Neither. Each is scored 0 to 5 and each contributes equally, alongside your converted Build a Sentence score, to the averaged Writing band.
Not on their own. If you covered every required point and the message is understandable overall, the descriptors put you at 3 or above. Errors have to actually obscure meaning to push a complete response below that.
No. The email targets roughly 100 to 150 words and the discussion asks for at least 100. Beyond that, extra words mostly add extra errors and eat the checking time that protects your framing and completion marks.
Blank, off topic, written in a language other than English, or copied verbatim from the prompt. A response that is nothing but one opening sentence before a cutoff also falls here, because none of the scoring dimensions can be assessed.
Your Build a Sentence total converts to a score out of 5 and is averaged with your email and discussion scores. That average maps to a Writing band from 1.0 to 6.0 in half-point steps, aligned to CEFR levels.
Report it. Many items legitimately have more than one grammatical arrangement, and those alternatives are stored on the item so the scorer credits any of them. A correct English ordering that was marked wrong is a scoring bug worth fixing, not a grammar disagreement.