Scoring

Introducing Prepamigo Scoring Engine 2.0 for Speaking

Sep 6, 2026

Prepamigo Scoring Engine 2.0 vs. frontier models as graders

Share of speaking responses where the score matched the independently verified score. 48 responses the systems had never seen, evenly spread across lower, middle and top scores. Each model was given the transcript and asked, in plain words, to score it on the CELPIP scale; average of three runs.

81%

Prepamigo Scoring Engine 2.0Prepamigo

62%

Claude Opus 5Anthropic

58%

GeminiGoogle

45%

Claude Sonnet 5Anthropic

45%

GPT-4o miniOpenAI

37%

GPT 5.5OpenAI

35%

GPT 5.6 solOpenAI

31%

GPT 5.6 lunaOpenAI

Prepamigo vs. the top chatbot plans

Canadian dollars. Prepamigo prices include tax. Claude Max (from US$100) and ChatGPT Pro (US$200) converted at 1.38, before tax. Accuracy is the share of held-out responses scored at the verified level when the named model is asked, in plain words, to score the response; average of three runs.

Prepamigo
Claude MaxOpus 5
ChatGPT ProGPT 5.6 sol
Monthly price
CA$24.98 (inc. tax)
CA$138+ (plus tax)
CA$276 (plus tax)
Scoring
Unlimited
Capped
Capped
Ready-made question bank
Yes
No
No
Practice sets
80
None
None
Practice tasks
2,000+
None
None
CELPIP format
Yes
No
No
Accuracy
81%
62%
35%
Top-level responses
71%
50%
25%

Our most accurate speaking grader yet. On responses it had never seen, Prepamigo Scoring Engine 2.0 matched verified scores 81% of the time, ahead of every frontier model we tested on the same job, and a clear step up from Engine 1.0, the grader it replaces.

Today we are releasing Prepamigo Scoring Engine 2.0 for Speaking, the system that grades every spoken practice response on Prepamigo. Engine 2.0 is a complete redesign of how we score. In place of a single grading pass, it uses two independent graders that compare each response directly against calibrated examples, vote several times, and reconcile their answers before a score is shown.

The result is a grader that is far more accurate than any frontier model used the way a student would use it. On a held-out set of real spoken responses with independently verified scores, Engine 2.0 landed on the correct score 81% of the time. Asked in plain words to score the same transcripts, the strongest of them, Claude Opus 5, reached 62%. Gemini reached 58%. Claude Sonnet 5 and GPT-4o mini reached 45%, GPT 5.5 37%, GPT 5.6 sol 35%, and GPT 5.6 luna 31%. When we then handed every model Engine 2.0’s own procedure, with real calibration examples for every task, the best of them climbed to 77% and still finished behind.

Accuracy matters most where it is hardest to achieve. Engine 2.0 recognises a top-level response 71% of the time. Asked plainly, no other model recognised more than 56%, and the GPT models recognised between 13% and 27%. Engine 2.0 also resists the most common failure of automated graders: scores that drift low and toward the middle of the scale no matter how strong or weak the response actually is. That drift was the weakness we saw most often in Engine 1.0.

Engine 2.0 is faster, more consistent, and produces feedback that quotes the student’s own words. It is live today for every Prepamigo user. This page explains what it does, how it compares, and what is new in the feedback that comes with every score.

Accuracy against verified scores

The question we care about is simple. When a student records a response, does the score on the screen match the score a trusted grader would have given? To answer it we built a test set of real spoken responses across all eight speaking task types, each with a score that had been independently verified, and we split the set evenly across lower, middle and top scores so that no single level could dominate the result.

The comparison uses 48 of those responses that were held back from the start. Nobody on the team used them while building or tuning Engine 2.0. Every response was transcribed word for word. Each model was then told it was an experienced CELPIP examiner, given the task prompt and the transcript, and asked to score the response from 0 to 12, three times, on different days. We averaged the three runs and counted how often the score matched the verified score’s level.

51%

fewer mistakes than Claude Opus 5

19 wrong scores per 100 responses, versus 38.

71%

fewer mistakes than GPT 5.6 sol

19 wrong scores per 100, versus 65.

71%

of top-level responses recognised

The best of the other models recognises 56%; GPT 5.6 sol recognises 25%.

Three things stand out in these numbers. First, the gap is not small. Nineteen points against the strongest frontier model and forty-six against the model behind ChatGPT’s paid plans, on a 48-response test, is the difference between nine and twenty-two extra correct scores. Second, the gap is not an artefact of a friendly test. The 48 responses were chosen before Engine 2.0’s design was final and were never touched during development. Third, the comparison is the one that matters to a student: it measures what you get when you paste a transcript into a chatbot and ask, which is how nearly everyone uses these models.

A word on what “correct” means here. Verified scores on this scale come in levels rather than single points, so we count a score as correct when it lands inside the verified level. A response verified at the top level, for example, is scored correctly if the engine gives it a 9, 10 or 11. Under a stricter reading that only accepts 10 and above, every system loses a few points and the ranking does not change.

SystemPlain requestGiven our full procedure
Prepamigo Scoring Engine 2.081.3%
Claude Opus 561.8%77.1%
Gemini58.3%70.8%
Claude Sonnet 545.1%68.8%
GPT-4o mini45.1%50.0%
GPT 5.536.8%68.8%
GPT 5.6 sol35.4%70.8%
GPT 5.6 luna30.6%62.5%

“Plain request” is a three-run average of the plain test. “Given our full procedure” is a single pass per response with the same transcripts, instructions and calibration examples as Engine 2.0’s own graders, plus a calibration note tuned to each model; Gemini also received the audio in that test. It is the best result each model could achieve in our hands, and it is not something a student can reproduce without the calibration examples.

Performance at every score level

An average accuracy figure can hide a great deal. A grader that is excellent on weak responses and hopeless on strong ones might post a respectable number while failing the students who most need a precise answer. So we report accuracy separately for each of the three levels in the test set: lower scores of 4 or 5, middle scores of 7 or 8, and top scores of 10 and above.

Most graders, human or machine, drift toward the middle. A strong response gets an 8 instead of a 10. A weak one gets a 6 instead of a 5. With nothing to compare against, the frontier models add a second habit: they drift downward, so that a weak response is often scored a 3 and a strong one a 7. Engine 2.0 was built specifically to resist both pulls, and the results show it most clearly at the top of the scale.

Lower scores · verified score of 4 or 5

16 held-out responses per system, plain request, average of three runs.

88%

Prepamigo Engine 2.0

77%

Claude Opus 5

75%

GPT-4o mini

56%

Gemini

54%

GPT 5.5

54%

GPT 5.6 sol

52%

GPT 5.6 luna

44%

Claude Sonnet 5

Middle scores · verified score of 7 or 8

16 held-out responses per system, plain request, average of three runs.

85%

Prepamigo Engine 2.0

63%

Gemini

63%

Claude Sonnet 5

58%

Claude Opus 5

44%

GPT-4o mini

29%

GPT 5.5

27%

GPT 5.6 sol

27%

GPT 5.6 luna

Top scores · verified score of 10 or above

16 held-out responses per system, plain request, average of three runs. Nearly every miss is a 7 or 8.

71%

Prepamigo Engine 2.0

56%

Gemini

50%

Claude Opus 5

29%

Claude Sonnet 5

27%

GPT 5.5

25%

GPT 5.6 sol

17%

GPT-4o mini

13%

GPT 5.6 luna

For students who are already strong, this is the difference that matters. Asked plainly, none of the other models recognises more than 56% of top-level responses; GPT 5.6 sol recognises one in four, and GPT 5.6 luna one in eight. Engine 2.0 recognises 71%.

At the lower level, Engine 2.0 leads at 88%, with Claude Opus 5 at 77% and GPT-4o mini at 75%. Every other model is around half, and Claude Sonnet 5 is at 44%. The miss is almost always a 3: the model sees a weak response and scores it below the level rather than in it. GPT-4o mini’s respectable figure here is the first sign of its problem, which is that it scores nearly everything low; at the top of the scale it recognises one response in six.

At the middle level, Engine 2.0 leads at 85%. Gemini and Claude Sonnet 5 are at 63%, Claude Opus 5 at 58%, and then a cliff: GPT-4o mini at 44% and the three larger GPT models at 27% to 29%. Middle-level responses contain a mixture of strengths and weaknesses that has to be weighed rather than spotted, and with nothing to weigh against, the GPT models see the weaknesses and hand out a 5 or 6.

At the top level, the gap is widest. Engine 2.0 correctly identifies 71% of top-level responses. Gemini identifies 56% and Opus 5 exactly half. Five of the seven other models hand a 7 or 8 to at least seven out of every ten responses that deserve a 10. This is the middle-drift problem in its purest form. A top-level response is not flawless; it contains the occasional slip, a restart, a plain sentence among the complex ones, and a grader measuring it against an imagined perfect answer deducts for each one. Engine 2.0 is calibrated against examples that show what a real top-level response looks like, slips included, and it is explicitly designed to compare against those examples rather than against perfection.

What if you give them our full procedure?

The plain-request numbers are not a verdict on how smart these models are. They are a verdict on what happens when a model is asked for a number with nothing to compare against. So we ran a second test, giving every model exactly what Engine 2.0’s own graders receive: the transcript, the task, two real calibration examples at each level for that specific task, and an instruction to name the nearest example for each of the four dimensions rather than to produce a number. Each model also got a short calibration note tuned to its own tendencies.

Given our full procedure: share of responses scored at the verified level

Same 48 held-out responses. Same transcripts, instructions and calibration examples for every system; each model graded every response once.

81%

Prepamigo Engine 2.0

77%

Claude Opus 5

71%

GPT 5.6 sol

71%

Gemini

69%

GPT 5.5

69%

Claude Sonnet 5

63%

GPT 5.6 luna

50%

GPT-4o mini

Every model improves, some dramatically. GPT 5.6 sol doubles from 35% to 71%. Opus 5 climbs from 62% to 77%. GPT-4o mini barely moves, from 45% to 50%, because it scores everything low regardless of what it is shown. And every model still finishes behind Engine 2.0. At the top of the scale the gap persists: Opus 5, GPT 5.6 sol, Gemini and GPT 5.5 all stop at 63% of top-level responses, Sonnet 5 at 38%, GPT-4o mini at zero, against 71% for Engine 2.0.

This second test tells you where the value in Engine 2.0 actually sits. It is not a smarter model; it uses models from this same list. It is real examples of what each level sounds like on each specific task, a question that forces the model to compare rather than guess, and then three more things a single pass cannot provide: two independent graders from different providers, three votes from each with the middle vote kept, and a reconciliation rule that takes the higher score when the two graders disagree strongly, because analysis showed the lower verdict was almost always the wrong one on strong responses.

Consistency and stability

A grader you can trust has to be consistent. If the same recording can receive an 8 today and a 10 tomorrow, neither number means very much, and a student cannot tell whether their practice is working. So alongside accuracy, we measure whether Engine 2.0 gives the same answer when asked the same question twice.

We ran Engine 2.0 over the 48 held-out responses three separate times, on different days, with nothing changed in between. It scored 81.3% on every run. Not approximately 81%: the same 39 correct responses out of 48 each time, give or take one. When we look at individual responses rather than the aggregate, 87.5% received an identical score across all three runs, and only 4.2% ever moved by two points or more between runs. Those few are responses that sit right on the boundary between two levels, where a small difference in judgement legitimately tips the score one way or the other.

The plain-request test on the frontier models moved far more between runs. GPT 5.6 sol scored 31%, 40% and 35% on its three runs; Opus 5 scored 62%, 65% and 58%. Ask a chatbot to grade the same transcript on two different days and you will, fairly often, get two different scores.

Consistency comes from two design decisions. Each grader in Engine 2.0 votes three times on every response and the middle vote is kept, which removes the occasional outlier a single pass would produce. And the two graders’ verdicts are reconciled by a fixed rule rather than by another model, so the same pair of verdicts always produces the same final score. Both decisions were tested directly: with a single vote instead of three, run-to-run agreement fell from 87.5% to 77.1%, and the share of responses jumping two or more points more than doubled.

How Engine 2.0 grades a response

Engine 2.0 is not a single model. It is a procedure, and the procedure is what makes the difference. Here is what happens in the ten or so seconds between a student pressing stop and a score appearing on the screen.

  1. The recording is transcribed word for word. Every filler, every restart, every self-correction is kept. This turned out to be essential. When we tested a “cleaned” transcript with the hesitations removed, accuracy fell by fifteen points, because the hesitations are part of the evidence a grader needs to judge how a response sounds.
  2. Two independent graders read the transcript. They are built on different underlying models from different providers, so they do not share the same blind spots. Each grader receives the task prompt, the transcript, and a small set of calibration examples for that exact task: real responses at the lower, middle and top level, two per level, so that each level is shown as a range rather than a single point.
  3. Each grader compares rather than scores. This is the central change from Engine 1.0. Instead of being asked for a number, each grader is asked, for each of four dimensions of the response, which calibration example it most closely resembles. There is no “somewhere in between” option. Having to commit to a nearest example is what stops the drift toward the middle. The score is then derived from the chosen level by a fixed rule, not by the model.
  4. Each grader votes three times. The middle vote on each dimension is kept. This removes the occasional off-day answer without averaging away genuine judgement.
  5. The two verdicts are reconciled. If the graders agree or differ by a single point, the final score is their average. If they differ by two or more, the higher score is used. That rule is not generosity; it is calibration. When we analysed every case where the two graders disagreed strongly, the lower verdict was almost always the wrong one, because the disagreement nearly always arose on a top-level response that one grader had been too cautious about.
  6. Safeguards are applied. A short set of fixed rules handles cases no grader should have to guess about: a response that is silent, a few words long, in another language, or entirely off topic receives a floor score regardless of what the graders return. In testing, silence scored a 3, a few-word response a 4, and an off-topic or non-English response a 5, with no failures across the set.

Two properties of the final design deserve emphasis. Engine 2.0 grades from the transcript alone, so it does not need the audio after transcription and does not depend on any one provider’s audio model. And because the two graders are from different providers, if one is unavailable the other continues, with a third model standing in so that a score is always produced.

What is new in your feedback

Engine 2.0 ships with a rebuilt feedback layer. It reads the two graders' evidence, not just the score, and it was tested against three kinds of student: someone meeting CELPIP for the first time, someone with weak English who wants the trick, and someone with half an hour a day and a test next week. Here is what changed.

  • Your own words, quoted back. Every point of feedback cites the exact phrase from your transcript that the graders reacted to. If a phrase is in quotation marks, it is copied verbatim; in our audit 157 of 158 quotes were. The feedback never invents something you did not say.
  • One thing to fix first. Each response gets a single top move: the one change that would raise your score the most, written in the plainest words on the page. Two more concrete actions follow it, then a three-item checklist to run through in your head before you start speaking.
  • Advice that fits the task. Giving advice, describing a scene and persuading someone are scored on different things. The feedback now knows what each of the eight task types rewards and speaks to that, instead of repeating the same generic tips across tasks.
  • Which dimension to practise first. The feedback tells you which of the four dimensions is your weakest and which is your strongest, so you know where the next point is, without hiding that behind a number.
  • What the task asked for and you missed. For task fulfillment, the feedback lists the specific parts of the prompt you did not cover, or says plainly that you covered them all.
  • Things you can actually rehearse. Every how-to is something you can say out loud before your next attempt. No advice to speak more clearly or reduce your accent: accent is not what listenability measures, and the feedback never comments on it.
  • No blame for the timer. A response cut off by the clock after it has completed the task is not marked down for the cut-off, and the feedback does not scold you for it.

When an independent judging model compared this feedback with what our previous system produced for the same responses, without knowing which was which, it preferred the new feedback in 20 of 24 comparisons, both judging as an examiner would and judging as a student would.

Unlimited practice for CA$25 a month

Accuracy is only useful if you can afford to use it every day. A speaking score you can trust is most valuable on your fortieth practice response, not your fourth, because that is when the score starts telling you whether a change in how you speak is actually working. So we priced Engine 2.0 to be used without counting.

Every Prepamigo plan includes unlimited scoring by Engine 2.0, unlimited attempts, and the full question bank: 80 complete practice sets and more than 2,000 practice tasks across Speaking, Writing, Reading and Listening, written to follow the task types, timing and format of the CELPIP General test. You press record, you get a score and evidence-based feedback in about ten seconds, and you can do it again as many times as you like.

Prepamigo

CA$25/ month

CA$24.98 month to month · CA$8.25 a month on the annual plan (CA$98.98 a year) · tax included · cancel any time

  • Unlimited scoring by Scoring Engine 2.0, with no daily, session or weekly cap.
  • 80 practice sets and 2,000+ tasks across all four sections, modelled on the real test format, so you never have to write your own prompts.
  • Record and go. Transcription, grading and feedback are handled for you; there is nothing to paste and nothing to prompt.
  • Feedback in your own words, with every quote traceable to what you actually said.

Claude Max, for comparison

US$100/ month

about CA$138 before tax · 20x tier US$200 a month · Pro US$20 with far tighter caps

  • Even Max is capped. A rolling five-hour usage window and a weekly limit, five or twenty times what Pro allows; when you hit either, you wait.
  • No question bank. You bring your own tasks, decide for yourself whether they resemble the real test, and keep track of what you have already practised.
  • You do the setup. Recording, transcribing, pasting the transcript and writing the grading instructions are all on you, every time.
  • A plain request, not a calibrated grader. Asked in plain words, Claude Opus 5 matched verified scores 62% of the time and Claude Sonnet 5 45%, against 81% for Engine 2.0.

Claude plan prices and usage terms are as published on claude.com in September 2026 and may change. CELPIP is a trademark of its owner; Prepamigo is an independent practice platform and is not affiliated with, endorsed by or connected to the test maker. Prepamigo practice tasks are original material written to follow the public test format.

Availability

Prepamigo Scoring Engine 2.0 for Speaking is live now for all Prepamigo users on every speaking task type. There is nothing to enable. Every new speaking response is scored by Engine 2.0, and every score comes with feedback drawn from the student’s own words.

A typical response is scored in about ten seconds, faster than Engine 1.0. Engine 2.0 also costs us less to run per response than the design it replaced, which is what allows us to run two graders with three votes each for every student without changing what Prepamigo costs.

We will publish updates to the figures on this page as the real-world validation described above completes. If you have a recording you believe Engine 2.0 scored wrongly, we want to see it: those cases are the fastest way we know to find the next improvement.

For the head-to-head comparisons, read Prepamigo vs Claude and Prepamigo vs ChatGPT. Or record a response and watch Engine 2.0 run.

Introducing Prepamigo Scoring Engine 2.0 for Speaking | PrepAmigo