Every TOEFL Speaking response — whether on Listen and Repeat or Take an Interview — is scored on a 0-to-5 scale against a small set of consistent criteria: how directly you address what's being asked, how fluently you speak, how clearly you pronounce your words, and how accurately and flexibly you use grammar and vocabulary. Understanding exactly what each score level requires is one of the highest-leverage things you can do to raise your Speaking band, because it turns vague advice like "speak more fluently" into a specific, checkable target.
This guide breaks down the rubric criteria for both current Speaking task types in detail. If you haven't yet reviewed the section's overall structure, start with our TOEFL Speaking section format guide first.
Take the free TOEFL Benchmark Test and get your estimated band score in about 45 minutes.
Regardless of task type, TOEFL Speaking responses are evaluated against four consistent dimensions. Understanding each one separately — rather than thinking of "good speaking" as one vague, unified quality — makes it far easier to diagnose exactly what's holding your score back.
A response doesn't need to be flawless across all four to score well — but a response that's notably weak on any single criterion, even if strong on the others, will be capped by that weak point. This is exactly why a response with perfect grammar but long silences and frequent filler words scores lower than test-takers often expect.
Each individual Speaking question is scored on the 0-to-5 scale described throughout this guide, and your overall Speaking section band — on the 1.0–6.0 scale described in our good TOEFL score guide — reflects your performance in aggregate across all Speaking questions, not any single response in isolation. This means one noticeably weaker response, caused by a single unfamiliar topic or momentary nerves, won't define your entire Speaking band the way it might feel like it does in the moment — but a consistent pattern of weakness on the same criterion across multiple responses will. This is exactly why diagnosing which specific criterion is weak, rather than reacting to any single response, is the more productive way to direct your Speaking preparation.
Take an Interview asks you to respond to a spoken interview question with a developed, spontaneous answer. Here's what each score level actually requires.
| Score | What It Requires |
|---|---|
| 5 — Fully Successful | Fully addresses the question; on-topic and well elaborated with clear reasons and details; conversational pace with natural pausing; pronunciation easily intelligible with supportive rhythm/intonation; accurate grammar and vocabulary with real range. |
| 4 — Generally Successful | Answers the question, on-topic and elaborated but may feel simple or less connected; may lack strong transitions; generally good pace with some pausing that slightly affects flow; mostly clear pronunciation with occasional effort needed; adequate grammar/vocabulary for general meaning. |
| 3 — Partially Successful | Addresses the question but clarity or elaboration is limited; generally on-topic but short or underdeveloped; frequent or lengthy pauses and filler words; intelligibility sometimes affected by pronunciation or rhythm; limited grammar/vocabulary range restricts precise meaning. |
| 2 — Mostly Unsuccessful | Only minimally connected to the question, sometimes just echoing question words; little or no relevant elaboration; meaning often difficult to understand; very limited grammar/vocabulary. |
| 1 — Unsuccessful | Only vaguely connected to the question; mostly unintelligible; mainly isolated words or phrases rather than full ideas. |
| 0 — No Credit | No response, entirely unintelligible, no English spoken, or completely unrelated content. |
Notice how much separates a 3 from a 5: it's not one dramatic flaw, but a combination of shorter elaboration, more frequent pausing, and a narrower vocabulary range — each individually modest, but compounding into a meaningfully lower score. This is why targeted, specific practice on each criterion tends to outperform vague "just talk more" advice.
Take the free TOEFL Benchmark Test and get your estimated band score in about 45 minutes.
Listen and Repeat asks you to listen to a sentence and repeat it back as accurately as possible. Unlike Take an Interview, this task isn't about developing your own ideas — it's scored primarily on how precisely you reproduce the original sentence, both in content and in delivery.
The single most important distinction for this task: content words (nouns, verbs, key adjectives) matter more than function words (small connecting words like "the," "a," "of"). Missing or changing a content word costs you significantly more than dropping a minor function word, so if you only catch part of a sentence, prioritize reproducing the content words accurately over forcing a complete but inaccurate sentence. See our Listen and Repeat task guide for the full technique.
Directly answer the question first, in your opening sentence, rather than building up to your answer gradually — this immediately establishes relevance. Then add at least one specific reason and one concrete detail or example, rather than stopping at a general statement. A response that states an opinion but never explains why will be capped by weak elaboration regardless of how fluent or grammatically accurate it is.
Practice speaking continuously for the task's full time limit without stopping, even if you have to fill a gap with a natural transition phrase rather than silence. Recording yourself and counting filler words and pause length is one of the most effective diagnostic exercises — most test-takers are surprised by how much more frequent these are in a recording than they perceived while speaking in the moment.
Focus on word-level stress and sentence-level rhythm and intonation, not just individual sound accuracy — these carry meaning as much as correct phonemes do, and are often the actual source of an intelligibility gap rather than mispronounced individual sounds. Recording and comparing your intonation against a native or highly fluent speaker's recording of the same content is a direct, practical way to identify specific gaps.
Deliberately vary your sentence structures rather than relying on the same simple pattern repeatedly — mixing in complex sentences with subordinate clauses, when used accurately, signals stronger control than a string of short, simple sentences even if each individual sentence is grammatically correct. Similarly, build a working vocabulary of precise, topic-relevant words rather than relying on the same small set of generic adjectives and verbs across every response.
Seeing the rubric applied side by side to two responses on the same prompt makes the criteria far more concrete than reading them in the abstract. Imagine an interview question asking whether it's better to study alone or in a group.
"I think, um, studying alone is better. Because, um, you can focus. Group is, um, distracting sometimes." This response is on-topic but underdeveloped — one reason given ("you can focus") without a specific example or further explanation. Frequent filler words disrupt fluency, and the vocabulary and sentence structures are simple and repetitive. This lands around a 3: partially successful, addressing the question but with limited elaboration and noticeable fluency issues.
"I personally prefer studying alone, mainly because it lets me control my own pace. When I study with a group, I often end up waiting for others to catch up on a concept I already understand, or feeling rushed if everyone else has moved ahead. Studying alone means I can spend extra time on whatever's actually difficult for me specifically, without worrying about anyone else's schedule." This response directly answers the question in the first sentence, develops it with a specific, concrete reason and a clear example, maintains a natural conversational pace without noticeable filler words, and uses varied sentence structure (a subordinate clause, a contrast, a qualifying phrase) rather than short, repetitive sentences. This is the kind of response the rubric describes at the 5 level.
Notice that the second response isn't longer because it rambles — it's longer because it's genuinely more developed, with concrete reasoning a listener can follow. That distinction, not raw word count, is what the elaboration criterion is actually measuring.
Start structured TOEFL preparation built around the current 2026 format.
Test-takers who don't know the rubric tend to prepare in a generic way — "practice speaking more," "try to sound confident" — which is directionally reasonable but too vague to target specific, fixable gaps efficiently. Test-takers who understand the four criteria can instead diagnose their own practice responses the way an examiner or AI grader would: is this response actually underdeveloped, or is it developed but delivered with too much hesitation? Is a low score coming from unclear pronunciation, or from a genuinely narrow vocabulary range that happens to sound clear?
This distinction matters because the fix for each gap is different. A relevance and elaboration gap responds to practicing response structure and having more ready reasoning frameworks. A fluency gap responds to timed speaking practice and filler-word awareness. A pronunciation gap responds to targeted listening-and-repeating drills on specific sounds or intonation patterns. Treating all low scores as the same generic "needs more speaking practice" problem wastes time on the wrong fix.
Once you understand the four criteria, use them to structure your own practice review rather than just recording a generic overall impression of each response. After each practice response, rate yourself explicitly on each of the four dimensions — relevance/elaboration, fluency, pronunciation, and grammar/vocabulary — rather than a single overall score. This surfaces exactly which criterion is holding your response back, the same way a detailed AI-graded or coach-reviewed response would, and lets you target your next practice session at that specific gap rather than practicing generically. See our sample answers for Take an Interview to calibrate what a genuinely strong response sounds like against each criterion.
Get AI-graded feedback on your Speaking responses against the full scoring rubric.