Your CELPIP Speaking result is a single level from 0 to 12, decided by three to five trained raters who score all eight of your tasks against four rated dimensions. There is no percentage, no raw mark, and no separate score for each task. All eight feed one Speaking level, and from level 3 upward that level equals the same CLB level.
That structure explains most of what confuses people about Speaking results. You cannot work out your level by counting anything, no single task carries its own published weight, and one stumble will not sink you, because several people score you independently and their scores are checked against each other.
#How is CELPIP Speaking scored?
By human raters, not by a computer. CELPIP describes the process in its published test-results material, and four details matter to you as a test taker.
The phrase to hold on to is tangible evidence. Raters do not award a level for a general impression of how confident you sounded. They assign a level in each dimension by finding evidence in your performance that matches the written descriptors for that level. That is why a vague, safe answer often scores lower than learners expect: there is less in it for a rater to point at.
#What are the four rated dimensions?
CELPIP publishes the four dimensions and the factors inside each one. These are the actual categories your speaking is judged on.
Two things in that list surprise almost everyone.
Grammar is not its own dimension. It sits inside Listenability, next to rhythm, pronunciation and pausing. That grouping tells you how CELPIP is thinking about it: grammar matters here because broken sentence structure makes you harder to follow, not because a rater is hunting for errors. The same logic covers self-correction. Stopping to restart a sentence is a Listenability event.
Tone and length sit under Task Fulfillment, not under style. Speaking to a close friend the way you would speak to a manager is not a small stylistic slip. It is a Task Fulfillment problem, in the same category as answering a different question from the one you were asked.
If you have read our guide to CELPIP Writing scoring, three of these four will look familiar. The difference is the third one: Writing is judged on Readability, Speaking on Listenability. Same idea, different channel.
#Who actually scores your speaking?
Real people, with published minimum qualifications. CELPIP lists them, and they are stricter than most test takers assume.
| Requirement | What CELPIP asks for |
|---|---|
| English proficiency | A native speaker of English, or a non-native speaker at CLB 11 or 12 |
| Education | A minimum of an undergraduate degree |
| Teaching and assessment | An ESL teaching certification recognised by TESL Canada, or graduate training in language education or linguistics, plus at least three years of relevant experience |
| Residency | Resident in Canada at the time of scoring |
Raters also certify before they start, and are monitored after. CELPIP runs rater performance analysis monthly, gives raters feedback on how closely they agree with other raters, and ends the contract of a rater who does not improve to standard.
The practical takeaway is reassuring, and it is worth holding on to if you are nervous about your accent. Your recording is not being judged by one person who may or may not like how you sound. It is being scored by several qualified raters who cannot see each other's decisions, whose agreement with each other is measured every month, and who are trained specifically to reduce the bias that human judgement introduces.
#What happens if the raters disagree?
A benchmark rater is brought in automatically. When the ratings for one performance are complete, they are inspected for agreement, and if they disagree the system assigns an additional rater. Benchmark raters are experienced raters who have shown consistent accuracy and reliability, and they assess the performance without knowing the ratings already given.
This is the part of CELPIP Speaking scoring that almost nobody explains, and it should change how you think about an answer you are unsure about. A response that sits between two levels does not get rounded down by whoever happened to hear it first. Disagreement triggers more scrutiny, not less.
#How does a set of dimension ratings become a level from 0 to 12?
Through a process CELPIP calls standard setting. Your component score is derived from the dimensional ratings the raters assigned, and that score is then transformed into a CELPIP level using rules established by English language experts.
Those experts worked with testing professionals to identify what a learner needs to be able to do at each performance level, analysed the test in detail, and decided what a test taker has to demonstrate to reach each CELPIP level. It is the reason a Speaking 9 is meant to mean the same thing in March as it does in September, and on your test form as on somebody else's.
Try a Task 1 Giving Advice question →
#Why is there no raw score or percentage for Speaking?
Because a raw count would not mean the same thing on two different test forms. CELPIP does not report raw scores for any component, and explains why: test forms are built to the same guidelines but can still vary slightly in difficulty, so a raw score of 30 would not carry the same meaning across forms. Scores are corrected for those differences and then reported as a level.
For Speaking the point goes further, because there is nothing countable to begin with. There is no set of right answers to total. There is a performance, and four dimensions of judgement applied to it by several people.
#Where do learners actually lose marks?
Vocabulary, more than any other dimension. We looked at more than 275,400 anonymised, aggregated Speaking responses from over 1,700 learners on HelloCelpip, and grouped every issue our evaluator flagged by the dimension it belongs to. About 32% of everything flagged is a Vocabulary issue. The other three dimensions sit close together well behind it, within about two points of each other — close enough that ranking them among themselves would be reading more into the numbers than they support.
Two honest caveats before you act on that. These are our evaluator's judgements, not a rater's, and they describe practice responses on HelloCelpip rather than official test performances. And the flat bottom three are the finding as much as the tall bar is: there is no single second problem to fix.
It is also worth noticing how this differs from Writing. On the Writing side of the same evaluation, the largest share of flagged issues is Readability, not Vocabulary. The two skills leak marks in different places, which is a good argument against studying them as one block.
#The same leak shows up on every task
We expected the picture to change from task to task, because the eight tasks ask for very different things. It does not.
We ran the same breakdown separately on three tasks chosen to be as unalike as possible — Task 1, where you give someone advice; Task 3, where you describe a picture you are looking at; and Task 6, where you deal with a difficult situation. Vocabulary came out as the largest share of flagged issues on all three, in a narrow band of roughly 31% to 33%, with Listenability behind it around 22% to 24% each time.
That is a more useful result than a ranking of hard tasks would have been. It says the vocabulary problem is not something one task exposes and the others hide, so there is no single task to drill as a shortcut. Whatever is causing it travels with you across the whole component, which is also why it is worth working on directly rather than hoping a stronger task will carry the average.
#What that means for how you practise
The useful thing about Vocabulary leading is that it is the dimension that responds fastest to deliberate work, and the least dependent on how nervous you feel on the day.
It is also the one people misread. A Vocabulary issue is usually not a small word count. It is the wrong word for the situation, or a phrase that is grammatically fine but not how anyone actually says it. Our own error analysis puts the overwhelming majority of vocabulary problems in exactly that category — collocation and precision — rather than in not knowing enough words. Memorising a themed list of advanced vocabulary targets a small slice of the real problem.
What works better is narrower: collect the phrases that fit the eight situations CELPIP actually asks about, and practise them in full sentences until they are automatic. Giving advice, describing a scene, and dealing with a difficult situation each have their own natural phrasing, and it is that phrasing raters are listening for.
The CELPIP Speaking practice test walks through all eight tasks if you want to see where each one puts its pressure.
How is CELPIP Speaking scored?
How many people score my CELPIP Speaking test?
Does my accent affect my CELPIP Speaking score?
Does grammar count in CELPIP Speaking?
Is one Speaking task worth more than the others?
What is a good CELPIP Speaking score?
Can I ask for my Speaking to be re-marked?
Confirm the four rated dimensions, the rating procedure, rater qualifications and re-evaluation rules on the official CELPIP Results page, and the level to CLB mapping on the CELPIP Score Comparison Chart. Immigration requirements are set per ability and change, so check them on canada.ca. Your HelloCelpip level uses the same 0 to 12 CELPIP scale, so you can track real progress between now and test day. Your official score comes from Paragon Testing Enterprises on test day. HelloCelpip is an independent study resource and is not affiliated with CELPIP or Paragon Testing Enterprises.
Official sources: CELPIP test results · CELPIP score comparison chart · IRCC language testing
Last updated: 10 August 2026