Share of responses scored at the verified level
48 held-out speaking responses, 16 per band. Each GPT model was given the transcript and asked, in plain words, to score it on the CELPIP scale; average of three runs. Engine 2.0 ran in its normal production configuration.
81%
Prepamigo Engine 2.0
45%
GPT-4o mini
37%
GPT 5.5
35%
GPT 5.6 sol
31%
GPT 5.6 luna
Prepamigo vs. the top chatbot plans
Canadian dollars. Prepamigo prices include tax. Claude Max (from US$100) and ChatGPT Pro (US$200) converted at 1.38, before tax. Accuracy is the share of held-out responses scored at the verified level when the named model is asked, in plain words, to score the response; average of three runs.
ChatGPT is probably the most common tool CELPIP candidates use to grade their own speaking practice. It is on everyone’s phone, it answers in seconds, and it will happily tell you what level your response was. The question this article answers is whether the level it tells you is the level you would actually get. We ran the test the way a student would: paste the transcript, ask for a score, and count.
The short answer: on 48 speaking responses none of the systems had seen, each with an independently verified score, Prepamigo Scoring Engine 2.0 landed on the correct level 81% of the time. Asked in plain words, GPT 5.6 sol, the model behind ChatGPT’s paid plans, landed on the correct level 35% of the time. GPT 5.5 managed 37%, GPT 5.6 luna 31%, and GPT-4o mini 45%, a figure that looks better than it is, because that model scores almost everything low and so happens to get the weak responses right.
We then did something most comparisons skip. We gave each GPT model our own grading procedure, with real calibration examples for every task and a forced comparison against them, to see how good it could possibly be. GPT 5.6 sol climbed to 71%, GPT 5.5 to 69%, luna to 63% and GPT-4o mini to 50%. All four still finished behind Engine 2.0, and all four still struggled most on top-level responses. The rest of this article walks through both tests, the results at each score level, why a general-purpose assistant drifts toward the middle of the scale, what the daily workflow looks like on each side, and what each option costs if you intend to practise seriously.
The plain test: what a student actually gets
Eighty-one percent against thirty-five. In a 48-response test, that is twenty-two extra correct scores for Engine 2.0 over GPT 5.6 sol. Put another way, Engine 2.0 makes 71% fewer scoring mistakes than GPT 5.6 sol, 70% fewer than GPT 5.5, 73% fewer than GPT 5.6 luna and 66% fewer than GPT-4o mini.
Three separate runs of the plain test gave GPT 5.6 sol 31%, 40% and 35%, and GPT 5.5 35%, 42% and 33%. That spread between runs is itself a finding: ask ChatGPT to grade the same transcript on two different days and you will, fairly often, get two different scores. Engine 2.0 scored 81.3% on each of its three runs.
Where the gap opens: score level by score level
An overall figure hides where the errors fall. A grader that is excellent on weak responses and hopeless on strong ones is fine for a beginner and actively harmful for someone chasing a 10. So here are the plain-test results split by verified band.
Lower-level responses (verified score 4 or 5)
16 held-out responses, plain prompt, average of three runs. Most misses are a score of 3.
88%
Prepamigo Engine 2.0
75%
GPT-4o mini
54%
GPT 5.5
54%
GPT 5.6 sol
52%
GPT 5.6 luna
On weak responses, Engine 2.0 leads at 88%. GPT-4o mini follows at 75%, which is the one place its habit of scoring low works in its favour. The three larger GPT models are at 52% to 54%: roughly half the time they score a 4 or 5 response a 3, which is the wrong band even though the direction is understandable. A beginner using ChatGPT to grade will be told they are worse than they are about half the time.
Middle-level responses (verified score 7 or 8)
16 held-out responses, plain prompt, average of three runs. Misses are almost all a 5 or 6.
85%
Prepamigo Engine 2.0
44%
GPT-4o mini
29%
GPT 5.5
27%
GPT 5.6 sol
27%
GPT 5.6 luna
The middle band is where the plain-prompt GPT models fall apart. Engine 2.0 is at 85%; GPT-4o mini at 44%; GPT 5.5, GPT 5.6 sol and luna between 27% and 29%. Middle-level responses contain a real mixture of strengths and weaknesses that has to be weighed, and with no calibration examples to weigh against, the GPT models see the weaknesses and hand out a 5 or 6. A 7 or 8 speaker asking GPT 5.6 sol for a score will hear the right band fewer than three times in ten.
Top-level responses (verified score 10 or above)
16 held-out responses, plain prompt, average of three runs. Nearly every miss is a 7 or 8.
71%
Prepamigo Engine 2.0
27%
GPT 5.5
25%
GPT 5.6 sol
17%
GPT-4o mini
13%
GPT 5.6 luna
At the top, Engine 2.0 recognises 71% of top-level responses. GPT 5.5 recognises 27%, GPT 5.6 sol 25%, GPT-4o mini 17%, and GPT 5.6 luna 13%. Every GPT model hands a 7 or 8 to at least seven out of every ten responses that deserve a 10.
Consider what that means for a strong candidate. You record eight responses, transcribe them, paste them into ChatGPT and ask for a score, and it tells you six of them are 7s and 8s. You conclude you are a solid middle-band speaker with work to do. You were at a 10 the whole time. That is not a hypothetical. It is what every GPT model did on most of the top-level responses in our test.
What if you give ChatGPT our full procedure?
It is tempting to read the plain-test numbers as GPT being a weak family of models. It is not. The problem is not intelligence. It is that a model asked for a number, with nothing to compare against, guesses, and when it guesses it guesses low and toward the middle.
So we ran the second test. Each GPT model received the same package Engine 2.0’s graders receive: the transcript, the task, two real calibration examples at each level for that exact task, and an instruction to name the nearest example for each of the four dimensions rather than to produce a number. Each model also got a short calibration note tuned to its own tendencies.
Given our full procedure: share of responses scored at the verified level
Same 48 held-out responses. Each GPT model graded every response once with the same transcripts, instructions and calibration examples as Engine 2.0's own graders.
81%
Prepamigo Engine 2.0
71%
GPT 5.6 sol
69%
GPT 5.5
63%
GPT 5.6 luna
50%
GPT-4o mini
The calibration examples make an enormous difference. GPT 5.6 sol doubles, from 35% to 71%. GPT 5.5 goes from 37% to 69%, luna from 31% to 63%. GPT-4o mini barely moves, from 45% to 50%, because it keeps scoring everything low regardless of what it is shown; with the full procedure it still recognised no top-level responses at all. This tells you where the value in a purpose-built grader actually sits: not in a smarter model, but in real examples of what each level sounds like on each specific task, and in a question that forces the model to compare rather than guess.
Even so, every GPT model stays behind. With the full procedure, GPT 5.6 sol is ten points behind Engine 2.0 overall, ten points behind in the middle band (75% against 85%), and eight points behind at the top (63% against 71%). Its remaining habit is over-rewarding polish: a competent 8 with smooth delivery gets pushed to a 9 or 10. GPT 5.6 luna does the opposite, pulling middle responses down toward 5, and finishes the middle band at 44%. A single model making a single pass still has a characteristic bias, and Engine 2.0 closes the remaining gap with two independent graders from different providers, three votes from each with the middle kept, and a reconciliation rule that takes the higher score when the two graders disagree strongly.
We also tried the obvious shortcut: keeping the plain prompt and adding instructions such as “do not drift toward the middle” or “a level 9 response need not be flawless.” They produced no measurable change. What works is changing the question, and giving the model something real to compare against.
The workflow gap: record and go, or do it all yourself
Accuracy numbers assume the grading actually happens. With ChatGPT, a surprising amount of work sits between pressing stop on your recording and reading a score, and some of that work changes the score.
First, you need a transcript. ChatGPT’s voice mode is a conversation, not a grader; to grade a recorded response you need text. That means either typing out what you said, which changes it, or running the recording through a transcription tool. And the transcript has to be verbatim. Every filler, every restart, every self-correction has to stay in. When we tested a “cleaned” transcript with the hesitations removed, accuracy fell by fifteen points, because those hesitations are exactly the evidence a grader needs to judge how a response sounds. Most consumer transcription tools, including the one built into most phones, clean by default.
Second, you need the task. ChatGPT can invent a CELPIP-style prompt if you ask, but it is guessing at the format, the timing and the kind of scenario the real test uses. You will not know whether the prompt it wrote resembles what you will face.
Third, if you want anything better than the plain-test numbers, you need calibration examples: real responses at each level for the same task type, with verified levels. That is the input that doubled GPT 5.6 sol’s accuracy in our test, and it is the one a student cannot produce at home.
Fourth, you do all of that again for the next response, and each one counts against your message allowance.
On Prepamigo, you pick a task from the bank, the timer runs, you speak, and about ten seconds after you stop, the score and feedback are on the screen. Transcription is verbatim by design, the calibration examples are attached to the task automatically, and there is nothing to paste and nothing to prompt. The eleventh response of the evening costs you the same effort as the first.
Consistency: the same response, the same score
A score you cannot reproduce is not a measurement. If the same recording gets an 8 on Monday and a 10 on Tuesday, you have learned nothing about your speaking. So we also tested repeatability.
Engine 2.0 was run over the 48 held-out responses three separate times on different days. It scored 81.3% every time. Looking at individual responses, 87.5% received an identical score across all three runs, and only 4.2% ever moved by two points or more, all of them responses sitting on the boundary between two bands.
The plain test on GPT 5.6 sol moved nine points between its best and worst run. That stability gap comes from voting. Each grader in Engine 2.0 votes three times on every response and the middle vote is kept. When we tested the same configuration with a single vote instead of three, identical scores across runs fell from 87.5% to 77%, and the share of responses jumping two or more points more than doubled. A single ChatGPT conversation is a single vote. We also checked whether setting a fixed random seed through the API made GPT models more repeatable. It did not; on GPT 5.6 luna, repeatability actually fell.
Feedback you can act on
ChatGPT is a good explainer. If you ask why it gave a score, it will tell you at length, in clear English, and it will suggest improvements. For a candidate who wants to talk through a response, that is genuinely valuable.
The risk is that fluent explanation is not the same as accurate diagnosis. A model that scored a 10 as a 7 will explain, convincingly, why it is a 7. It will find something to point to. And because it did not have to commit to evidence before scoring, the explanation is written after the fact to justify a number it has already chosen.
Engine 2.0 works the other way round. Each grader has to point to evidence for every dimension before it names a level, and the feedback is written from that evidence. It quotes your own words: in an audit of 158 quotations across a sample of feedback, 157 appeared verbatim in the transcript. When an independent judging model compared Engine 2.0’s feedback with our previous system’s feedback on the same responses, blind, it preferred Engine 2.0 in 20 of 24 comparisons, both judging as an examiner would and judging as a student would.
We also checked that the feedback lands criticism in the right place, by taking strong responses and deliberately damaging one dimension at a time. In three of four cases the damaged dimension was the one whose score fell the most. That is a check a chatbot cannot easily pass, because it has no fixed dimensions to be checked against.
What it costs to practise every day
A speaking score is most useful on your fortieth response, not your fourth. That is when it starts telling you whether a change in how you speak is actually working. So the practical question is what each tool costs to use in volume.
ChatGPT Plus is US$20 a month, about CA$28, billed monthly, as published in September 2026. It gives you the GPT 5.6 family with a cap on messages per rolling window; long transcripts with calibration examples attached are exactly the kind of message that uses a lot of that allowance, and once you hit it you wait or drop to a smaller model. ChatGPT Pro, at US$200 a month, about CA$276 before tax, raises the limits substantially. The free tier routes much of its traffic to smaller models, and on this test the smallest of them recognised almost no top-level responses. None of the plans include a question bank, a timer, verbatim transcription, calibration examples, or anything specific to CELPIP.
Prepamigo is CA$24.98 a month, or CA$8.25 a month on the annual plan, which is CA$98.98 for the year, tax included, and you can cancel at any time. Every plan includes unlimited scoring by Engine 2.0 with no daily, session or weekly cap, unlimited attempts, and the full question bank: 80 complete practice sets and more than 2,000 practice tasks across Speaking, Writing, Reading and Listening, written to follow the task types, timing and format of the CELPIP General test.
Put simply: for less than the price of ChatGPT Plus, you get a grader that is more than twice as accurate on a plain comparison, that never asks you to wait, and that comes with the questions and calibration examples already in place.
When ChatGPT is the better tool
This is not an argument for never opening ChatGPT while preparing for CELPIP. A serious candidate might well use both.
- Talking through a response. If you want to understand why a reason was weak or brainstorm three better ones, a conversation is the right shape. Engine 2.0 gives you a verdict and evidence; it does not chat.
- Vocabulary and phrasing. Asking for five more natural ways to say something is quick and useful.
- Speaking to something that talks back. ChatGPT’s voice mode is a reasonable way to get comfortable speaking English out loud. It is not a grader, but it is company.
- Writing practice. Written responses do not need transcription, so the workflow gap is smaller. GPT accuracy on writing was not part of this test and we make no claim about it.
- Anything outside the test. ChatGPT is a general assistant and Prepamigo is not.
What ChatGPT should not be, on the evidence here, is the thing that tells you whether you are ready. For that you want a grader that was built for the job, tested on responses it had not seen, and measured level by level.
What is new in your feedback
Engine 2.0 ships with a rebuilt feedback layer. It reads the two graders' evidence, not just the score, and it was tested against three kinds of student: someone meeting CELPIP for the first time, someone with weak English who wants the trick, and someone with half an hour a day and a test next week. Here is what changed.
- Your own words, quoted back. Every point of feedback cites the exact phrase from your transcript that the graders reacted to. If a phrase is in quotation marks, it is copied verbatim; in our audit 157 of 158 quotes were. The feedback never invents something you did not say.
- One thing to fix first. Each response gets a single top move: the one change that would raise your score the most, written in the plainest words on the page. Two more concrete actions follow it, then a three-item checklist to run through in your head before you start speaking.
- Advice that fits the task. Giving advice, describing a scene and persuading someone are scored on different things. The feedback now knows what each of the eight task types rewards and speaks to that, instead of repeating the same generic tips across tasks.
- Which dimension to practise first. The feedback tells you which of the four dimensions is your weakest and which is your strongest, so you know where the next point is, without hiding that behind a number.
- What the task asked for and you missed. For task fulfillment, the feedback lists the specific parts of the prompt you did not cover, or says plainly that you covered them all.
- Things you can actually rehearse. Every how-to is something you can say out loud before your next attempt. No advice to speak more clearly or reduce your accent: accent is not what listenability measures, and the feedback never comments on it.
- No blame for the timer. A response cut off by the clock after it has completed the task is not marked down for the cut-off, and the feedback does not scold you for it.
When an independent judging model compared this feedback with what our previous system produced for the same responses, without knowing which was which, it preferred the new feedback in 20 of 24 comparisons, both judging as an examiner would and judging as a student would.
FAQs
Can ChatGPT score my CELPIP speaking? It will give you a number. Asked plainly on 48 unseen responses with verified scores, GPT 5.6 sol was right 35% of the time, GPT 5.5 37%, luna 31%, GPT-4o mini 45%, against 81% for Engine 2.0. On top-level responses GPT 5.6 sol was right 25% of the time.
Is Prepamigo more accurate than ChatGPT? 81% against 35% for GPT 5.6 sol with a plain request. Given our full procedure with calibration examples, GPT 5.6 sol reached 71%, still ten points behind.
How much does each cost? ChatGPT Pro is US$200 a month, about CA$276 before tax; Plus is US$20 with message caps. Prepamigo is CA$24.98 a month tax included, or CA$8.25 a month on the annual plan, with unlimited scoring and 80 practice sets.
Which GPT model should I use if I do use ChatGPT? With a plain prompt, none of them does well. With calibration examples, GPT 5.6 sol. Never GPT-4o mini if you are anywhere near a 10.
For the same comparison against Claude, read Prepamigo vs Claude. For the full leaderboard of eight systems, read which AI scores CELPIP speaking most accurately. Or skip the transcription and record a response.
