Product

Prepamigo vs Claude for CELPIP Speaking: Which One Gives You the Right Score?

Sep 7, 2026

Share of responses scored at the verified level

48 held-out speaking responses, 16 at each of three levels. Claude was given the transcript and asked, in plain words, to score it on the CELPIP scale; average of three runs. Engine 2.0 ran in its normal production configuration.

81%

Prepamigo Engine 2.0

62%

Claude Opus 5

45%

Claude Sonnet 5

Prepamigo vs. the top chatbot plans

Canadian dollars. Prepamigo prices include tax. Claude Max (from US$100) and ChatGPT Pro (US$200) converted at 1.38, before tax. Accuracy is the share of held-out responses scored at the verified level when the named model is asked, in plain words, to score the response; average of three runs.

Prepamigo
Claude MaxOpus 5
Monthly price
CA$24.98 (inc. tax)
CA$138+ (plus tax)
Scoring
Unlimited
Capped
Ready-made question bank
Yes
No
Practice sets
80
None
Practice tasks
2,000+
None
CELPIP format
Yes
No
Accuracy
81%
62%
Top-level responses
71%
50%

If you have ever pasted a speaking transcript into Claude and asked it to grade you like a CELPIP rater, you already know the appeal. It answers instantly, it explains itself, and it never gets tired. The question is whether the number it gives you is the number you would actually get. We ran that test the way a student would run it: the same transcript, a plain request for a score, and then we counted.

The short answer: on 48 speaking responses the systems had never seen, each with an independently verified score, Prepamigo Scoring Engine 2.0 landed on the correct level 81% of the time. Asked in plain words to score the same transcripts, Claude Opus 5 landed on the correct level 62% of the time and Claude Sonnet 5 45% of the time. The gap is widest exactly where it matters most for anyone chasing a high score: on top-level responses, Engine 2.0 recognised 71%, Opus 5 recognised 50%, and Sonnet 5 recognised 29%.

We then did something most comparisons skip. We gave Claude our own grading procedure, with real calibration examples for every task and a forced comparison against them, to see how good it could possibly be. Opus 5 climbed to 77% and Sonnet 5 to 69%. Both still finished behind Engine 2.0, and both still struggled most on top-level responses. The rest of this article walks through both tests, what the numbers look like at each score level, why a general-purpose assistant drifts toward the middle of the scale, what the daily workflow looks like on each side, and what each option costs if you intend to practise seriously.

The plain test: what a student actually gets

Eighty-one percent against sixty-two and forty-five. In a 48-response test, the difference between Engine 2.0 and Opus 5 is nine extra correct scores, and the difference between Engine 2.0 and Sonnet 5 is seventeen. Put another way, Engine 2.0 makes 51% fewer scoring mistakes than Opus 5 and 66% fewer than Sonnet 5.

Three separate runs of the plain test gave Opus 5 62%, 65% and 58%, and Sonnet 5 44%, 46% and 46%. That spread of several points between runs is itself a finding: ask Claude to grade the same transcript on two different days and you will, fairly often, get two different scores. Engine 2.0 scored 81.3% on each of its three runs.

Where the gap opens: score level by score level

Most graders, human and machine, drift toward the middle of the scale. A strong response becomes an 8 instead of a 10; a weak one becomes a 6 instead of a 5. That drift is invisible in an average and obvious the moment you split results by level. With a plain prompt, Claude adds a second habit on top: it drifts downward, so that even weak responses are often scored below their level.

Lower-level responses (verified score 4 or 5)

16 held-out responses, plain prompt, average of three runs. Sonnet 5's misses here are mostly scores of 3.

88%

Prepamigo Engine 2.0

77%

Claude Opus 5

44%

Claude Sonnet 5

On weak responses, Engine 2.0 leads at 88%, Opus 5 follows at 77%, and Sonnet 5 falls to 44%. This is the easy end of the scale, and Opus 5 handles it reasonably. Sonnet 5 does not: it scores a 4 or 5 response a 3 more often than not, which is the wrong level even though the direction is understandable. A student at this level who uses Sonnet 5 to grade will be told they are worse than they are, most of the time.

Middle-level responses (verified score 7 or 8)

16 held-out responses, plain prompt, average of three runs. Misses are almost all a 6.

85%

Prepamigo Engine 2.0

63%

Claude Sonnet 5

58%

Claude Opus 5

In the middle, Engine 2.0 is at 85% and the two Claude models are at 63% and 58%. Middle-level responses contain a mixture of strengths and weaknesses that has to be weighed rather than spotted, and with no calibration examples to weigh against, both Claude models tend to see the weaknesses and hand out a 6. A 7 or 8 speaker will hear that they are a 6 roughly four times in ten.

Top-level responses (verified score 10 or above)

16 held-out responses, plain prompt, average of three runs. Nearly every miss is a 7 or 8.

71%

Prepamigo Engine 2.0

50%

Claude Opus 5

29%

Claude Sonnet 5

At the top, Engine 2.0 recognises 71% of top-level responses. Opus 5 recognises half of them. Sonnet 5 recognises 29%, which means it hands a 7 or 8 to seven out of every ten responses that deserve a 10.

Think about what that means in practice. Suppose your speaking really is at a 10 today. You record eight responses, transcribe them, paste them into Claude Sonnet 5 and ask for a score, and it tells you six of them are 7s and 8s. You conclude you are not ready. You spend three more weeks drilling. You were ready the whole time. That is not a hypothetical failure mode; it is the most common single error in our entire test, and it falls hardest on the candidates who have worked hardest.

If you use Claude to grade and it keeps giving you 7s and 8s, do not automatically believe it. In our test, a top-level response was more likely to get an 8 from Sonnet 5 than a 10. An 8 from Claude is a reason to get a second opinion, not a verdict.

What if you give Claude our full procedure?

It is tempting to read the plain-test numbers as Claude being a weak model. It is not. Opus 5 is, by some distance, the strongest general-purpose grader we have tested. The problem is not intelligence. It is that a model asked for a number, with nothing to compare against, guesses, and when it guesses it guesses low and toward the middle.

So we ran the second test. Claude received the same package Engine 2.0’s graders receive: the transcript, the task, two real calibration examples at each level for that exact task, and an instruction to name the nearest example for each of the four dimensions rather than to produce a number. Each model also got a short calibration note tuned to its own tendencies.

Given our full procedure: share of responses scored at the verified level

Same 48 held-out responses. Each Claude model graded every response once with the same transcripts, instructions and calibration examples as Engine 2.0's own graders.

81%

Prepamigo Engine 2.0

77%

Claude Opus 5

69%

Claude Sonnet 5

The calibration examples make a large difference. Opus 5 climbs from 62% to 77%; Sonnet 5 from 45% to 69%. Lower-level accuracy for both reaches 88%, the same as Engine 2.0. Middle-level accuracy reaches 81% for both. This is worth dwelling on, because it tells you where the value in a purpose-built grader actually sits: not in a smarter model, but in real examples of what each level sounds like on each specific task, and in a question that forces the model to compare rather than guess.

Even so, both models stay behind at the top of the scale. With the full procedure, Opus 5 recognises 63% of top-level responses and Sonnet 5 recognises 38%, against 71% for Engine 2.0. A single model making a single pass still holds back on a strong response with a slip in it. Engine 2.0 closes the remaining gap with three more things: two independent graders from different providers, three votes from each with the middle vote kept, and a reconciliation rule that takes the higher score when the two graders disagree strongly, because analysis showed the lower verdict was almost always the wrong one on strong responses.

We also tried the obvious shortcut: keeping the plain prompt and adding instructions such as “do not drift toward the middle” or “a level 9 response need not be flawless.” They produced no measurable change. The model agrees with the instruction and then does what it was going to do anyway. What works is changing the question, and giving it something real to compare against.

The workflow gap: record and go, or do it all yourself

Accuracy numbers assume the grading actually happens. With Claude, a surprising amount of work sits between pressing stop on your recording and reading a score.

First, you need a transcript. Claude’s chat interface does not grade audio recordings, so you need to produce text. That means either typing out what you said, which changes it, or running the recording through a transcription tool. And here is a detail that cost us a lot of accuracy before we understood it: the transcript has to be verbatim. Every filler, every restart, every self-correction has to stay in. When we tested a “cleaned” transcript with the hesitations removed, accuracy fell by fifteen points, because those hesitations are exactly the evidence a grader needs to judge how a response sounds. Most consumer transcription tools clean by default.

Second, you need the task. Claude does not have a bank of CELPIP-style prompts with the right timing and format. You either write your own, which means guessing at what the real test asks, or find one elsewhere and paste it in with the transcript.

Third, if you want anything better than the plain-test numbers, you need calibration examples: real responses at each level for the same task type, with verified levels. That is the input that took Opus 5 from 62% to 77% in our test, and it is the one a student cannot produce at home.

Fourth, you do all of that again for the next response. And the one after that.

On Prepamigo, you pick a task from the bank, the timer runs, you speak, and about ten seconds after you stop, the score and feedback are on the screen. Transcription is verbatim by design, the calibration examples are attached to the task automatically, and there is nothing to paste and nothing to prompt. The eleventh response of the evening costs you the same effort as the first.

Consistency: the same response, the same score

A score you cannot reproduce is not a measurement. If the same recording gets an 8 on Monday and a 10 on Tuesday, you have learned nothing about your speaking and something about the grader’s mood. So we also tested whether each system gives the same answer twice.

Engine 2.0 was run over the 48 held-out responses three separate times. It scored 81.3% on every run. Looking at individual responses rather than the aggregate, 87.5% received an identical score across all three runs, and only 4.2% ever moved by two points or more. Those few are responses that sit right on the boundary between two levels, where a small difference in judgement legitimately tips the score one way or the other.

The plain test on Claude moved several points between runs at the aggregate level, and more at the level of individual responses. That consistency gap is not an accident of the models involved. It comes from voting. Each grader in Engine 2.0 votes three times on every response and the middle vote is kept, which removes the occasional outlier a single pass produces. When we tested the same configuration with a single vote instead of three, run-to-run agreement fell from 87.5% to 77.1%. A single Claude conversation is a single vote.

Feedback you can act on

A score tells you where you are. Feedback tells you what to do next, and this is an area where Claude is genuinely good. It writes clearly, it is patient, and if you ask it why it gave a particular score it will explain at whatever length you like. For a candidate who wants to talk through a response, that is valuable.

The risk is the same as with the score: fluent explanation is not the same as accurate diagnosis. A model that scored a 10 as a 7 will explain, convincingly, why it is a 7. It will find something to point to. And because Claude did not have to commit to evidence before scoring, the explanation is written after the fact, to justify a number it has already chosen.

Engine 2.0 works the other way round. Because each grader has to point to evidence for every dimension it judges, the feedback layer receives that evidence directly: the exact phrases from your transcript that supported the verdict on content, vocabulary, how you sound, and whether you completed the task. The feedback is written from that evidence and quotes your own words. In our audit of 158 quotations across a sample of feedback, 157 appeared verbatim in the transcript. When we asked an independent judging model to compare Engine 2.0’s feedback with the feedback our previous system produced, without knowing which was which, it preferred Engine 2.0 in 20 of 24 comparisons, both when judging as an examiner would and when judging as a student would.

We also tested whether the feedback lands the criticism in the right place. We took strong responses and deliberately damaged one aspect of each: inserting grammar errors and fillers, swapping precise words for vague ones, stripping out supporting detail, or removing the core action the task asked for. In three of the four cases, the dimension we had damaged was the one whose score fell the most. That is the kind of check a chatbot cannot easily pass, because it has no fixed dimensions to be checked against.

What it costs to practise every day

A speaking score is most useful on your fortieth response, not your fourth. That is when it starts telling you whether a change in how you speak is actually working. So the practical question is not what each tool costs, but what it costs to use it in volume.

Claude Pro is US$20 a month, about CA$28, or US$17 a month on annual billing, as published on claude.com in September 2026. It comes with a rolling five-hour usage window and a weekly limit; when you hit either, you wait. Long transcripts with calibration examples attached are exactly the kind of message that uses a lot of that allowance. The Max plans, from US$100 a month, about CA$138 before tax, raise the limits substantially but do not remove them. None of the plans include a question bank, a timer, transcription, calibration examples, or anything specific to CELPIP.

Prepamigo is CA$24.98 a month, or CA$8.25 a month on the annual plan, which is CA$98.98 for the year, tax included, and you can cancel at any time. Every plan includes unlimited scoring by Engine 2.0 with no daily, session or weekly cap, unlimited attempts, and the full question bank: 80 complete practice sets and more than 2,000 practice tasks across Speaking, Writing, Reading and Listening, written to follow the task types, timing and format of the CELPIP General test.

Put simply: for less than the price of Claude Pro, you get a grader that is far more accurate on the responses that matter most, that never asks you to wait, and that comes with the questions and the calibration examples already in place.

When Claude is the better tool

This is not an article arguing that you should never open Claude while preparing for CELPIP. There are several things it does better than a specialised scoring engine, and a serious candidate might well use both.

  • Talking through a response. If you want to understand why a particular reason was weak, or brainstorm three better ones, a conversation is the right shape. Engine 2.0 gives you a verdict and evidence; it does not chat.
  • Vocabulary and phrasing help. Asking Claude for five more natural ways to say something is quick and useful, and its suggestions are generally good.
  • Writing practice. Written responses do not need transcription, so the workflow gap is much smaller. Claude’s accuracy on writing was not part of this test, and we make no claim about it here.
  • Anything outside the test. Immigration paperwork questions, study schedules, explaining a grammar point: Claude is a general assistant and Prepamigo is not.

What Claude should not be, on the evidence here, is the thing that tells you whether you are ready. For that you want a grader that was built for the job, tested on responses it had not seen, and measured level by level.

What is new in your feedback

Engine 2.0 ships with a rebuilt feedback layer. It reads the two graders' evidence, not just the score, and it was tested against three kinds of student: someone meeting CELPIP for the first time, someone with weak English who wants the trick, and someone with half an hour a day and a test next week. Here is what changed.

  • Your own words, quoted back. Every point of feedback cites the exact phrase from your transcript that the graders reacted to. If a phrase is in quotation marks, it is copied verbatim; in our audit 157 of 158 quotes were. The feedback never invents something you did not say.
  • One thing to fix first. Each response gets a single top move: the one change that would raise your score the most, written in the plainest words on the page. Two more concrete actions follow it, then a three-item checklist to run through in your head before you start speaking.
  • Advice that fits the task. Giving advice, describing a scene and persuading someone are scored on different things. The feedback now knows what each of the eight task types rewards and speaks to that, instead of repeating the same generic tips across tasks.
  • Which dimension to practise first. The feedback tells you which of the four dimensions is your weakest and which is your strongest, so you know where the next point is, without hiding that behind a number.
  • What the task asked for and you missed. For task fulfillment, the feedback lists the specific parts of the prompt you did not cover, or says plainly that you covered them all.
  • Things you can actually rehearse. Every how-to is something you can say out loud before your next attempt. No advice to speak more clearly or reduce your accent: accent is not what listenability measures, and the feedback never comments on it.
  • No blame for the timer. A response cut off by the clock after it has completed the task is not marked down for the cut-off, and the feedback does not scold you for it.

When an independent judging model compared this feedback with what our previous system produced for the same responses, without knowing which was which, it preferred the new feedback in 20 of 24 comparisons, both judging as an examiner would and judging as a student would.

FAQs

Can Claude score my CELPIP speaking practice? It will give you a number. Asked plainly on 48 unseen responses with verified scores, Claude Opus 5 landed on the verified level 62% of the time and Claude Sonnet 5 45% of the time, against 81% for Prepamigo Scoring Engine 2.0. On top-level responses Sonnet 5 was right 29% of the time.

Is Prepamigo more accurate than Claude? Yes: 81% against 62% and 45% with a plain request. Given our full procedure with calibration examples, Opus 5 reached 77% and Sonnet 5 69%, still behind, and still weakest at the top of the scale.

How much does each cost? Claude Max starts at US$100 a month, about CA$138 before tax, with usage caps; Claude Pro is US$20 with tighter caps. Prepamigo is CA$24.98 a month tax included, or CA$8.25 a month on the annual plan, with unlimited scoring and 80 practice sets.

Why does Claude keep giving my strong responses a 7 or 8? A model asked for a number with nothing to compare against drifts low and toward the middle, and a top-level response is never flawless. Asked plainly, Sonnet 5 gave a top-level response a 9 or higher only 29% of the time.

To see how the scoring actually works, read inside Scoring Engine 2.0. To see the full leaderboard including GPT and Gemini, read which AI scores CELPIP speaking most accurately. Or skip the reading and record a response.

Prepamigo vs Claude for CELPIP Speaking: Which One Gives You the Right Score? | PrepAmigo