Share of responses scored at the verified level
48 held-out responses, 16 per band. Each model was given the transcript and asked, in plain words, to score it on the CELPIP scale; average of three runs. Engine 2.0 ran in its normal production configuration.
81%
Prepamigo Engine 2.0
62%
Claude Opus 5
58%
Gemini
45%
Claude Sonnet 5
45%
GPT-4o mini
37%
GPT 5.5
35%
GPT 5.6 sol
31%
GPT 5.6 luna
Prepamigo vs. the top chatbot plans
Canadian dollars. Prepamigo prices include tax. Claude Max (from US$100) and ChatGPT Pro (US$200) converted at 1.38, before tax. Accuracy is the share of held-out responses scored at the verified level when the named model is asked, in plain words, to score the response; average of three runs.
Plenty of CELPIP candidates now grade their own speaking practice with an AI chatbot. Record a response, transcribe it, paste it in, ask for a score. It is fast and free or nearly free, and the explanations sound authoritative. What nobody tells you is how often the number is right, or which model to trust, or whether the way you ask changes the answer. We wanted to know, so we built the test and ran it on eight systems, two ways.
The first way is the chart above: paste the transcript, ask for a score, the way a student does. Among general-purpose models, Claude Opus 5 was the most accurate at 62%, Gemini followed at 58%, and everything else was below half: Claude Sonnet 5 and GPT-4o mini at 45%, GPT 5.5 at 37%, GPT 5.6 sol at 35%, GPT 5.6 luna at 31%. A purpose-built system, Prepamigo Scoring Engine 2.0, reached 81% on the same responses.
The second way is the one we think is most instructive: we gave every model our own grading procedure, with real calibration examples and a forced comparison, to see how good each could possibly be. Every model improved, some of them dramatically, and every model still finished behind Engine 2.0. We should be upfront that Prepamigo is our product. We have tried to make both tests fair to everyone else, and where a model beat us on a particular level, we say so.
The plain-prompt leaderboard
Three tiers are visible in the chart at the top. Engine 2.0 at 81% stands alone. Claude Opus 5 at 62% and Gemini at 58% form a second tier: usable, if you accept being wrong two times in five. Everything else is below 50%, which is to say worse than a coin flip between the three bands would be if you knew nothing about the response at all.
That is the defining feature of the plain test: every general-purpose model grades harshly. The average score they gave a verified top-level response ranged from 7.2 (luna) to 8.6 (Opus 5). The average they gave a verified lower-level response ranged from 3.7 to 4.3, below the band for all but Opus 5 and Gemini. With no examples of what each level actually sounds like, the models compare the response against an imagined perfect one and deduct.
Results at each score level
An overall figure hides where the errors fall, and where they fall is what determines whether a grader is useful to you. Here are the plain-test results by band.
Lower-level responses (verified 4 or 5)
16 held-out responses per system, plain prompt, average of three runs.
88%
Prepamigo Engine 2.0
77%
Claude Opus 5
75%
GPT-4o mini
56%
Gemini
54%
GPT 5.5
54%
GPT 5.6 sol
52%
GPT 5.6 luna
44%
Claude Sonnet 5
Weak responses should be the easy end of the scale, and with a plain prompt they are not. Engine 2.0 is at 88%, Opus 5 at 77%, GPT-4o mini at 75%. Everything else is around half, and Sonnet 5 is at 44%. The miss is almost always a 3: the model sees a weak response and scores it below the band rather than in it. GPT-4o mini’s respectable figure here is the first sign of its problem, which is that it scores nearly everything low.
Middle-level responses (verified 7 or 8)
16 held-out responses per system, plain prompt, average of three runs.
85%
Prepamigo Engine 2.0
63%
Gemini
63%
Claude Sonnet 5
58%
Claude Opus 5
44%
GPT-4o mini
29%
GPT 5.5
27%
GPT 5.6 sol
27%
GPT 5.6 luna
The middle band separates the field. Engine 2.0 is at 85%. Gemini and Sonnet 5 are at 63%, Opus 5 at 58%. Then a cliff: GPT-4o mini at 44% and the three larger GPT models at 27% to 29%. Middle-level responses contain a real mixture of strengths and weaknesses that has to be weighed, and with nothing to weigh against, the GPT models see the weaknesses and hand out a 5 or 6. A 7 or 8 speaker asking GPT 5.6 sol for a score will hear the right band fewer than three times in ten.
Top-level responses (verified 10 or above)
16 held-out responses per system, plain prompt, average of three runs. Nearly every miss is a 7 or 8.
71%
Prepamigo Engine 2.0
56%
Gemini
50%
Claude Opus 5
29%
Claude Sonnet 5
27%
GPT 5.5
25%
GPT 5.6 sol
17%
GPT-4o mini
13%
GPT 5.6 luna
The top band is where the differences matter most, and where the plain prompt does the most damage. Engine 2.0 recognises 71% of top-level responses. Gemini recognises 56% and Opus 5 exactly half. Then Sonnet 5 at 29%, GPT 5.5 at 27%, GPT 5.6 sol at 25%, GPT-4o mini at 17% and GPT 5.6 luna at 13%. Five of the seven general-purpose models hand a 7 or 8 to at least seven out of every ten responses that deserve a 10.
If you are a strong candidate grading yourself with a chatbot, this is the number to remember. The chance that a plain request to GPT 5.6 sol will tell you a 10 is a 10 is one in four.
The second leaderboard: given our full procedure
The plain-test numbers are not a verdict on how smart these models are. They are a verdict on what happens when a model is asked for a number with nothing to compare against. So we ran the second test, giving every model the calibration examples and the forced comparison that Engine 2.0’s own graders use.
Given our full procedure: share of responses scored at the verified level
Same 48 held-out responses. Same transcripts, instructions and calibration examples for every system; each model graded every response once. Gemini also received the audio.
81%
Prepamigo Engine 2.0
77%
Claude Opus 5
71%
GPT 5.6 sol
71%
Gemini
69%
GPT 5.5
69%
Claude Sonnet 5
63%
GPT 5.6 luna
50%
GPT-4o mini
Every model improves. Opus 5 goes from 62% to 77%. GPT 5.6 sol doubles, from 35% to 71%. GPT 5.5 goes from 37% to 69%, Sonnet 5 from 45% to 69%, luna from 31% to 63%, Gemini from 58% to 71%. GPT-4o mini barely moves, from 45% to 50%, because it scores everything low regardless of what it is shown.
The order also changes. With a plain prompt, the Claude models were far ahead of the GPT models. With the full procedure, GPT 5.6 sol overtakes Sonnet 5 and ties Gemini, and four models cluster between 69% and 71%, statistically indistinguishable on this sample. Opus 5 stays on top of the general-purpose field at 77%, four points behind Engine 2.0.
At the level view, the full procedure lifts lower-band accuracy to 88% for four models, the same as Engine 2.0, and middle-band accuracy to 81% for the two Claude models. The top band is where the gap persists: Opus 5, GPT 5.6 sol, Gemini and GPT 5.5 all stop at 63%, Sonnet 5 at 38%, and GPT-4o mini at zero, against 71% for Engine 2.0. A single model making a single pass still holds back on a strong response with a slip in it. Engine 2.0 closes that remaining gap with two independent graders from different providers, three votes from each with the middle vote kept, and a reconciliation rule that takes the higher score when the two graders disagree strongly, because analysis showed the lower verdict was almost always the wrong one on strong responses.
Why the models drift low and toward the middle
The pattern is so consistent across models from three different companies that it cannot be a quirk of any one of them. It is a property of the task.
When a model is asked to place a response on a scale, it hedges toward the centre when it is unsure, because the centre is where a wrong answer costs least. And when it has no examples of what a level actually sounds like, it compares the response against an imagined perfect one and deducts for every slip. A top-level response is never flawless. It contains a restart, a plain sentence among the complex ones, a hesitation before a hard word. Measured against perfection, those slips cost a notch or two, and a 10 becomes an 8 or a 7.
Instructions do not fix this. We tried telling models explicitly that a level 9 does not require an error-free response, that under-scoring a strong response is as serious as over-scoring a weak one. Each instruction was accepted and then ignored in practice. When we fed a model the very calibration example it had been told represented a top-level response and asked it to score it, it gave it an 8.
What does fix it, mostly, is changing the question from “what number?” to “which example?” and supplying real examples. That is the difference between the two leaderboards. Even then, a single pass still drifts a little, because the calibration examples are not perfect and one reading of them is one reading. Two graders, three votes and a reconciliation rule are what close the rest.
Notes on each model
- Claude Opus 5. The best general-purpose grader on both tests: 62% plain (62%, 65% and 58% across three runs) and 77% with the full procedure. Its plain-prompt weakness is the middle band, where it hands out 6s; its full-procedure weakness is the top band, where it stops at 63%. It is also the most expensive model here by a wide margin.
- Gemini. Second on the plain test at 58%, identical across all three runs, and the most even across bands, with 56%, 63% and 56%. With the full procedure and the audio recording it reached 71%, level with GPT 5.6 sol, which worked from text alone. Hearing the recording did not help it beyond what a verbatim transcript provides.
- Claude Sonnet 5. 45% plain, with a peculiar profile: reasonable in the middle band at 63%, poor on weak responses at 44% because it scores them a 3, and poor at the top at 29%. With the full procedure it climbs to 69% but stays at 38% on top-level responses. If Sonnet 5 keeps giving you 8s, do not believe it.
- GPT-4o mini. 45% plain, 50% with the full procedure, and in both cases the number flatters it. It scores almost everything low: 75% on weak responses, 17% and then 0% on top-level ones. Whatever you do, do not grade your speaking with this model.
- GPT 5.5. 37% plain (35%, 42%, 33%), 69% with the full procedure. Its plain-prompt middle band is 29%. It was the text grader in an earlier version of our own pipeline and was replaced for weakness in exactly that band.
- GPT 5.6 sol. 35% plain (31%, 40%, 35%), the largest run-to-run swing of any model, and 71% with the full procedure, the biggest improvement of any model. A capable grader when shown examples, and a poor one when not. Its full-procedure habit is over-rewarding polish in the middle band.
- GPT 5.6 luna. Last on both tests: 31% plain and 63% with the full procedure. It pulls middle responses down toward 5 and recognised 13% of top-level responses plainly. Cheap enough to be useful as a second opinion inside a two-grader system; not accurate enough to use alone.
What a purpose-built system adds
Engine 2.0 is not a better model. It uses models from this same leaderboard. What it adds is procedure, and each piece of that procedure was chosen by measurement.
- Verbatim transcription, automatically. Every filler and restart kept, because removing them cut accuracy by fifteen points in testing. You do not have to find a transcription tool that does this; most consumer ones clean by default.
- Calibration examples attached to every task. Two real responses per level for each of the eight task types, already in place. This is the single input that doubled GPT 5.6 sol’s accuracy, and the one you cannot produce at home.
- A comparison, not a number. Each grader names the nearest example on each of four dimensions; the score follows by rule. This is what stops the drift.
- Two graders from two providers. Different models have different blind spots. Two of them, reconciled by a fixed rule, beat either alone.
- Three votes each, middle kept. This did not raise accuracy, but it took run-to-run consistency from 77% to 87.5%. A single chatbot answer is a single vote.
- Reconciliation that takes the higher score on strong disagreement. Because the lower verdict was almost always the wrong one on strong responses. A plain average fell to the mid-sixties in testing.
- Safeguards. Silence, a few words, off-topic and non-English responses receive floor scores by rule rather than by a model’s guess.
The result is 81% overall and 71% at the top, identical across three separate runs, in about ten seconds per response with nothing to paste.
If you are going to use a chatbot anyway
Some readers will keep grading with a chatbot, and that is a reasonable choice for a candidate on a tight budget or one who wants to talk through their responses. These steps are the difference between the two leaderboards, and they will get you closer to the second one than to the first.
- Use a verbatim transcript. Keep your fillers, restarts and self-corrections. If your transcription tool cleans them out, the grader is grading a better response than you gave.
- Give it examples, not a rubric. One or two real responses at each level for the same task type, with their levels labelled, do more than any description of what a level 9 looks like. If you do not have verified examples, this is the step you cannot replicate at home, and it is the one that matters most.
- Ask which example, not what number. For each of the four dimensions, ask the model to name the nearest example and forbid “in between.” Derive the score yourself from the levels it names.
- Ask more than once. Score the same response two or three times in fresh conversations and keep the middle answer. GPT 5.6 sol swung nine points between runs on the plain test.
- Pick a model with a top-band record. Opus 5 or Gemini if you are asking plainly; Opus 5 or GPT 5.6 sol if you have examples. Not GPT-4o mini, ever, if you are anywhere near a 10.
- Distrust a 7 or 8. On every general-purpose model, a 7 or 8 on a response you thought was excellent is more likely to be a miss than a verdict.
Which model to trust, depending on where you are
The right reading of these leaderboards depends on your current level, because the models fail in different places.
- If you are at a 4 or 5 today, Opus 5 asked plainly will tell you so about three times in four; most other models will call you a 3 about half the time. Your problem is not really the grader; it is the feedback, and for that you want something that quotes your words back to you rather than summarising.
- If you are at a 7 or 8, avoid the GPT models with a plain prompt, which will call you a 5 or 6 seven times in ten. Gemini and Sonnet 5 are the most reliable plain-prompt graders in the middle band at 63%. Engine 2.0 is at 85%.
- If you are at a 10 or above, this is where the choice matters most. Asked plainly, no general-purpose model will recognise a 10 more than 56% of the time, and the GPT models will do so between 13% and 27% of the time. Engine 2.0 recognises seven in ten and, more importantly, gives the same answer every time you ask.
The pattern across all three bands is the same one the whole article has been describing. The models are not stupid; they are cautious, and caution with nothing to compare against means a low number. The further you are from the middle, the more a grader’s caution costs you, and the more it matters that the grader was built to resist it.
What is new in your feedback
Engine 2.0 ships with a rebuilt feedback layer. It reads the two graders' evidence, not just the score, and it was tested against three kinds of student: someone meeting CELPIP for the first time, someone with weak English who wants the trick, and someone with half an hour a day and a test next week. Here is what changed.
- Your own words, quoted back. Every point of feedback cites the exact phrase from your transcript that the graders reacted to. If a phrase is in quotation marks, it is copied verbatim; in our audit 157 of 158 quotes were. The feedback never invents something you did not say.
- One thing to fix first. Each response gets a single top move: the one change that would raise your score the most, written in the plainest words on the page. Two more concrete actions follow it, then a three-item checklist to run through in your head before you start speaking.
- Advice that fits the task. Giving advice, describing a scene and persuading someone are scored on different things. The feedback now knows what each of the eight task types rewards and speaks to that, instead of repeating the same generic tips across tasks.
- Which dimension to practise first. The feedback tells you which of the four dimensions is your weakest and which is your strongest, so you know where the next point is, without hiding that behind a number.
- What the task asked for and you missed. For task fulfillment, the feedback lists the specific parts of the prompt you did not cover, or says plainly that you covered them all.
- Things you can actually rehearse. Every how-to is something you can say out loud before your next attempt. No advice to speak more clearly or reduce your accent: accent is not what listenability measures, and the feedback never comments on it.
- No blame for the timer. A response cut off by the clock after it has completed the task is not marked down for the cut-off, and the feedback does not scold you for it.
When an independent judging model compared this feedback with what our previous system produced for the same responses, without knowing which was which, it preferred the new feedback in 20 of 24 comparisons, both judging as an examiner would and judging as a student would.
FAQs
Which AI model is most accurate at scoring CELPIP speaking? Asked plainly: Claude Opus 5 at 62%, Gemini 58%, Claude Sonnet 5 and GPT-4o mini 45%, GPT 5.5 37%, GPT 5.6 sol 35%, GPT 5.6 luna 31%. Prepamigo Scoring Engine 2.0 reached 81% on the same responses.
ChatGPT or Claude? Claude, clearly, when asked plainly: 62% and 45% against 35% and 37%. With our full procedure the gap narrows to 77% for Opus 5 against 71% for GPT 5.6 sol.
Why do they all give my strong responses a 7 or 8? Models with nothing to compare against guess low and toward the middle, and a strong response always has small slips. Asked plainly, no general-purpose model recognised more than 56% of top-level responses.
How do I get a better score from a chatbot? Verbatim transcript, example responses at each level, ask which example rather than what number, ask more than once and keep the middle answer. That procedure doubled GPT 5.6 sol from 35% to 71% in our test.
For a closer look at the two head-to-heads, read Prepamigo vs Claude and Prepamigo vs ChatGPT. For how the procedure works step by step, read inside Scoring Engine 2.0. Or record a response and skip the transcription.
