Scoring

Is Prepamigo's CELPIP Speaking Scoring Accurate? The Numbers, Honestly

Sep 6, 2026

Engine 2.0 on 48 held-out responses

Share of responses scored inside the verified band. Three separate runs on different days each produced 81.3% overall.

81%

Exact band

Prepamigo vs. the top chatbot plans

Canadian dollars. Prepamigo prices include tax. Claude Max (from US$100) and ChatGPT Pro (US$200) converted at 1.38, before tax. Accuracy is the share of held-out responses scored at the verified level when the named model is asked, in plain words, to score the response; average of three runs.

Prepamigo
Claude MaxOpus 5
ChatGPT ProGPT 5.6 sol
Monthly price
CA$24.98 (inc. tax)
CA$138+ (plus tax)
CA$276 (plus tax)
Scoring
Unlimited
Capped
Capped
Ready-made question bank
Yes
No
No
Practice sets
80
None
None
Practice tasks
2,000+
None
None
CELPIP format
Yes
No
No
Accuracy
81%
62%
35%
Top-level responses
71%
50%
25%

It is the right question to ask of any practice tool, and most tools do not answer it. A speaking score is only worth something if it tells you what a real rater would say, and the only way to know whether it does is to measure. So this article is the measurement, laid out plainly, with the parts that flatter us and the parts that do not.

The answer in one paragraph: on 48 speaking responses that Prepamigo’s Scoring Engine 2.0 had never seen, each with an independently verified score, the engine landed on the verified level 81% of the time. It produced the same 81% on each of three separate runs. It is most accurate on weak and middle-level responses, at 88% and 85%, and least accurate on top-level responses, at 71%, where its characteristic mistake is to give a 10 an 8. Asked in plain words to score the same transcripts, the best general-purpose AI model we tested reached 62%, and the weakest reached 31%; given our full grading procedure, the best reached 77%.

What follows is how we got those numbers, what they mean at each score level, how stable they are, where the errors fall, how the score compares with the alternatives you might be using instead, what is new in the feedback, and how to read your own score in light of all of it. If you only want the practical advice, skip to the section on reading your score. If you want to decide whether to trust the number at all, read the whole thing.

What “accurate” means here

Before any number, the definition. We are not claiming that a Prepamigo score predicts your test result. No practice tool can promise that, and anyone who does is guessing. What we measure is narrower and more honest: given a spoken response whose score has been independently verified, how often does the engine produce a score at the same level?

CELPIP speaking is reported on a scale that runs to 12, and verified scores in our reference data come in three bands: lower, meaning 4 or 5; middle, meaning 7 or 8; and top, meaning 10 and above. We count the engine as correct when its score lands inside the verified band. A response verified at the top is scored correctly if the engine gives it a 9, 10 or 11. Under a stricter reading that only accepts 10 and above, every figure on this page drops by a few points and every comparison stays in the same order.

The headline numbers

What does 81% mean in practice? If you record ten responses at a steady level, the engine will put about eight of them in the right band. It will not put a 5 in the 10 band or a 10 in the 5 band; that did not happen once in 144 scored responses across the three runs. The mistakes it makes are the mistakes of a careful rater on a borderline response, and the next section shows where they concentrate.

Accuracy at each score level

Accuracy by verified band

16 held-out responses per band. Each bar is the share of that band's responses scored inside the band.

88%

Lower (4–5)

85%

Middle (7–8)

71%

Top (10+)

Lower-level responses: 88%. Weak responses are the easiest to recognise. They tend to be short, to reuse the prompt’s wording, and to leave parts of the task undone. The engine catches those signs most of the time. When it misses, it misses upward: in every case, the response was scored an 8 rather than a 4 or 5. The cause is almost always a response that sounds fluent and confident but says little; delivery earns it a level it has not earned on content. We know this because we can see which dimensions the graders placed high, and it is listenability every time.

Middle-level responses: 85%. These are the hardest to grade in one sense, because they contain a real mixture of strengths and weaknesses that has to be weighed rather than spotted. The engine’s misses here go in both directions. A few responses with visible hesitation but solid ideas were pulled down to a 5 or 6; a few polished-sounding responses were pushed up to a 10. The reconciliation step between the two graders catches most of these, which is why this figure is as high as it is.

Top-level responses: 71%. This is the weakest band and the one we talk about most, because the error is so consistent. When the engine misses a top-level response, it gives it an 8. Not a 7, not a 6; an 8, in every case but a handful of 9s that count as correct. The engine sees a strong response and holds back one notch.

This deserves context. The drift toward 8 on strong responses is not a Prepamigo quirk; it is the characteristic failure of every automated grader we have tested, human-calibrated or not. Our previous engine had it far worse. Asked plainly to score the same responses, the best general-purpose model recognised 56% of top-level responses; the model behind ChatGPT’s paid plans recognised 25%; and even given our full procedure, none of them beat 63%. Seventy-one percent is the best figure we have achieved, and it is still the number we most want to improve.

The 8 problem, from your side of the screen

The level view tells you how the engine does on a response of a known level. You are looking at it the other way round: you have a score and want to know how much to believe it. So here is the same data turned around. For each score the engine gives, how often was the response actually at that level?

  • If the engine says 4 or 5: the response was at the lower level 93% of the time. The remainder were middle-level responses with heavy hesitation. A low score from Engine 2.0 is almost always right.
  • If the engine says 10: the response was at the top level 92% of the time. A 10 from Engine 2.0 is a 10.
  • If the engine says 8: the response was at the middle level 67% of the time, at the top level 23% of the time, and at the lower level 10% of the time. An 8 is the one score worth reading carefully.

Put plainly: an 8 from Engine 2.0 means “middle, probably, but possibly better.” Roughly one response in four that receives an 8 was actually a 10. If you get an 8 on a response you thought was excellent, do not conclude you are not ready. Record the same task again. A consistent 10 across three attempts is a 10; a consistent 8 across three attempts is an 8; a mix is a response on the boundary, and the feedback will tell you which dimension is holding it there.

The most useful habit for a strong candidate is to treat an 8 as a question rather than a verdict. The feedback quotes the phrases the graders reacted to. If those phrases are all about hesitation and restarts, you are probably a 10 on content and a 9 on delivery, and one more take will show it.

Does the same recording get the same score?

Accuracy without consistency is luck. If the same recording could get an 8 today and a 10 tomorrow, the 81% figure would be an average over noise and you could not use the score to track anything. So we tested repeatability directly.

The 48 held-out responses were scored three separate times, on different days, with nothing changed in between. The overall figure was 81.3% on every run. At the level of individual responses, 87.5% received an identical score all three times, and only 4.2% ever moved by two points or more. Those few are responses that sit right on the boundary between two bands, where a small difference in judgement legitimately tips the score one way or the other.

The consistency comes from a specific design choice. Each of the engine’s two graders votes three times on every response, and the middle vote is kept. When we tested the same configuration with a single vote instead of three, identical scores across runs fell from 87.5% to 77%, and the share of responses jumping two or more points more than doubled. Voting did not change the accuracy figure. It changed whether the accuracy figure means anything.

For you, consistency means a change in your score is a change in your speaking. If you record the same task type three times over a week and the score moves from 8 to 10, that is not the grader having a good day. It is you.

How this compares with what you might use instead

An accuracy number means more with something to compare it to. So we ran the same 48 held-out responses through the general-purpose AI models people actually use to grade their practice, the way a person would: we gave each model the verbatim transcript and the task, told it it was an experienced CELPIP examiner, and asked for a score from 0 to 12. Three runs each, averaged.

Share of responses scored inside the verified band

Same 48 held-out responses. Each model was asked, in plain words, to score the transcript; average of three runs. Engine 2.0 ran in its normal production configuration.

81%

Prepamigo Engine 2.0

62%

Claude Opus 5

58%

Gemini

45%

Claude Sonnet 5

45%

GPT-4o mini

37%

GPT 5.5

35%

GPT 5.6 sol

31%

GPT 5.6 luna

Engine 2.0 leads every model by between nineteen and fifty-one points. Asked plainly, every general-purpose model grades harshly: the miss on a weak response is usually a 3, the miss on a strong one a 7 or 8. Claude Opus 5 and Gemini are the only two above half. The three larger GPT models are between 31% and 37%, and recognise a top-level response between 13% and 27% of the time.

We then gave every model our full procedure: the same calibration examples for the exact task, the same forced comparison, and a short calibration note tuned to its own tendencies. That is the best each model could achieve in our hands, and it is not something a student can reproduce without the examples. Opus 5 climbed to 77%, GPT 5.6 sol and Gemini to 71%, GPT 5.5 and Sonnet 5 to 69%, luna to 63% and GPT-4o mini to 50%. Every one of them still finished behind Engine 2.0, and none recognised more than 63% of top-level responses.

What is new in your feedback

Engine 2.0 ships with a rebuilt feedback layer. It reads the two graders' evidence, not just the score, and it was tested against three kinds of student: someone meeting CELPIP for the first time, someone with weak English who wants the trick, and someone with half an hour a day and a test next week. Here is what changed.

  • Your own words, quoted back. Every point of feedback cites the exact phrase from your transcript that the graders reacted to. If a phrase is in quotation marks, it is copied verbatim; in our audit 157 of 158 quotes were. The feedback never invents something you did not say.
  • One thing to fix first. Each response gets a single top move: the one change that would raise your score the most, written in the plainest words on the page. Two more concrete actions follow it, then a three-item checklist to run through in your head before you start speaking.
  • Advice that fits the task. Giving advice, describing a scene and persuading someone are scored on different things. The feedback now knows what each of the eight task types rewards and speaks to that, instead of repeating the same generic tips across tasks.
  • Which dimension to practise first. The feedback tells you which of the four dimensions is your weakest and which is your strongest, so you know where the next point is, without hiding that behind a number.
  • What the task asked for and you missed. For task fulfillment, the feedback lists the specific parts of the prompt you did not cover, or says plainly that you covered them all.
  • Things you can actually rehearse. Every how-to is something you can say out loud before your next attempt. No advice to speak more clearly or reduce your accent: accent is not what listenability measures, and the feedback never comments on it.
  • No blame for the timer. A response cut off by the clock after it has completed the task is not marked down for the cut-off, and the feedback does not scold you for it.

When an independent judging model compared this feedback with what our previous system produced for the same responses, without knowing which was which, it preferred the new feedback in 20 of 24 comparisons, both judging as an examiner would and judging as a student would.

How to read your score

  1. A 4 or 5 is almost certainly right. The engine identifies weak responses 88% of the time, and when it says 4 or 5 it is right 93% of the time. Read the feedback for which dimension is lowest and work on that one.
  2. A 10 is a 10. When the engine says 10, the response was at the top level 92% of the time. You are ready on that task type. Move to the ones you avoid.
  3. An 8 is a question. Two-thirds of the time it means middle. One time in four it means top. Record the same task again and look at the pattern across attempts.
  4. Trust changes more than levels. Because the score is stable, a change across attempts is a change in your speaking. Repeating the same task type is the fastest way to find out whether a new habit is working.
  5. Read the quotes. The phrases quoted in your feedback are what the graders reacted to. They are the specific words to change.

FAQs

How accurate is Prepamigo’s speaking score? On 48 unseen responses with verified scores: 81% at the exact band, identical across three runs. By band: 88% lower, 85% middle, 71% top.

Is it the same as my real score? No practice tool can promise that. What we measure is how often the engine agrees with a verified score on the same response. Read your score as a level with a direction, and look at the pattern across attempts.

Why did I get an 8 on a response I thought was excellent? That is the engine’s largest remaining error: one 8 in four was actually a 10. Record the same task again. A consistent 10 across attempts is a 10.

Will the same recording get the same score twice? 87.5% identical across three runs; only 4.2% moved by two or more points, all on band boundaries.

To see how the engine works step by step, read inside Scoring Engine 2.0. To see the same test run on GPT, Claude and Gemini, read which AI scores CELPIP speaking most accurately. Or record a response and see the numbers for yourself.

Is Prepamigo's CELPIP Speaking Scoring Accurate? The Numbers, Honestly | PrepAmigo