Prepamigo Writing Scoring Engine vs. frontier models as graders
Share of CELPIP writing responses where the score matched the independently verified score. 36 official essays, evenly spread across weak, middle and strong levels and both writing tasks. Each model was given the task and the essay and asked, in plain words, to score it on the CELPIP scale.
86%
Prepamigo Writing Scoring EnginePrepamigo
78%
GPT 5.6 solOpenAI
69%
GPT 5.6 lunaOpenAI
61%
GeminiGoogle
61%
Claude Opus 5Anthropic
58%
GPT-4o miniOpenAI
56%
Claude Sonnet 5Anthropic
Prepamigo vs. the top chatbot plans
Canadian dollars. Prepamigo prices include tax. Claude Max (from US$100) and ChatGPT Pro (US$200) converted at 1.38, before tax. Accuracy is the share of 36 official essays scored at the verified level when the named model is asked, in plain words, to score the essay; single run.
Prepamigo’s writing scorer matched verified scores on 86% of official CELPIP essays, ahead of every frontier model given the same job, and on the middle level, where most candidates are, it scored every essay correctly.
This page introduces the Prepamigo Writing Scoring Engine, the system that grades every email and survey response written on Prepamigo. It is not a general-purpose model asked to grade. It is a scorer built for one job: placing a CELPIP writing response at the level a trained rater would place it, and explaining that placement in the writer’s own words.
The result is a grader that is more accurate than any frontier model used the way a student would use it. On 36 official CELPIP writing responses with independently verified scores, the engine landed on the correct level 86% of the time. Asked in plain words to score the same essays, GPT 5.6 sol reached 78%. GPT 5.6 luna reached 69%, Gemini and Claude Opus 5 61%, GPT-4o mini 58%, and Claude Sonnet 5 56%.
Accuracy matters most where candidates actually are. A 7 or 8 is the most common real writing score, and it is the score that decides whether a target has been met. The engine placed all twelve middle-level essays inside that band. The best frontier model placed nine; Claude Opus 5 placed three. This page explains how the engine grades, how it compares, what is new in the feedback that comes with every score, and what it costs to practise without counting.
Accuracy against verified scores
The question we care about is simple. When a student submits an essay, does the score on the screen match the score a trained rater would have given? To answer it we used a set of official CELPIP writing responses, each with an independently verified score, covering both writing tasks and split evenly across weak, middle and strong levels: twelve essays per level, 36 in all. The six sample essays the engine uses for calibration are not among them.
Every system saw the same task, including the required points for the email task and the options for the survey task, and the same essay. Each frontier model was told it was an experienced CELPIP writing examiner and asked for one score from 0 to 12. The engine ran as it does in production. A score is correct when it lands inside the verified level: 4 or 5 for a weak essay, 7 or 8 for a middle one, 10 to 12 for a strong one.
37%
fewer mistakes than GPT 5.6 sol
5 wrong scores in 36, versus 8.
64%
fewer mistakes than Claude Opus 5
5 wrong scores in 36, versus 14.
100%
of middle-level essays scored correctly
The best of the other models: 75%. Claude Opus 5: 25%.
Three things stand out. First, the gap over the best frontier model is eight points, which on 36 essays is three extra correct scores, and it held on every repeat run. Second, the gap over the Claude models is 25 to 30 points, and it is concentrated in the middle of the scale, where a general-purpose model reads a clean, generic essay as a strong one. Third, the comparison is the one that matters to a student: it measures what you get when you paste an essay into a chatbot and ask, which is how nearly everyone uses these models.
| System | All essays | Weak | Middle | Strong |
|---|---|---|---|---|
| Prepamigo Writing Scoring Engine | 86% | 75% | 100% | 83% |
| GPT 5.6 sol | 78% | 83% | 75% | 75% |
| GPT 5.6 luna | 69% | 83% | 58% | 67% |
| Gemini | 61% | 58% | 42% | 83% |
| Claude Opus 5 | 61% | 67% | 25% | 92% |
| GPT-4o mini | 58% | 83% | 67% | 25% |
| Claude Sonnet 5 | 56% | 50% | 33% | 83% |
Share of essays scored inside the verified level, single run. Twelve essays per level, so one essay is worth about eight points at the level view and three points overall.
Consistency and stability
A grader you can trust has to give the same answer twice. After the headline run we scored all 36 essays twice more on different days. The engine scored 86%, 89% and 86%, with the same twelve middle-level essays inside the band every time and no essay moving by more than one point between runs. GPT 5.6 sol scored 78%, 83% and 81% on the same three days.
The engine runs the scoring model with reasoning switched off and at a fixed temperature, which is what makes the number repeatable. A single chatbot answer carries a few points of luck in either direction; a score that is stable across days is a score you can plan around.
Two tasks, two kinds of mistake
CELPIP writing has two tasks, and they fail in different ways. The email task gives you a situation and three or four points you must cover. The survey task gives you two options and asks you to pick one and defend it. A grader that treats them the same will be wrong about both.
On the email task the engine was right on 89% of essays. The mistakes that cost candidates most here are not grammar; they are a missing point, a tone that does not fit the reader, or an opening and closing that are not there. The rule layer checks each of those before the rater gives an opinion, and a missing point holds task fulfillment to 8 or below however polished the English is. A chatbot asked for a score in plain words does not know which points were required unless you paste them, and even then it rarely checks them one by one.
On the survey task the engine was right on 83% of essays. The survey task rewards a clear position and two developed reasons, and punishes a response that hedges. The engine reads the stance first: a response that never commits cannot score above the middle on task fulfillment. Chatbots tend to reward a balanced answer, which in a survey response is a mistake.
The point of a task-aware grader is that the score reflects what a rater would react to on that task, not a general impression of the English. A middle-level email is usually a middle-level email because of a thin third point; a middle-level survey response is usually one because a reason is asserted and not developed. The engine tells you which, by name.
Where the engine is not the best
The engine leads overall and in the middle. It is not the leader at both ends, and it is fairer to say so than to hide it in an average.
On weak essays, the three GPT models each scored 83% against the engine’s 75%. The engine’s misses were all a notch high, on short, tidy essays that did not develop anything. A GPT model, which reads generic writing as weak, gets these right. If you are at the weak level and only want confirmation of that, GPT 5.6 sol will give it to you slightly more often. It will also, three times in twelve, tell a middle-level writer they are a 6.
On strong essays, Claude Opus 5 scored 92% against the engine’s 83%. Claude reads polished writing generously, which on a strong essay is correct. The same generosity gave seven of twelve middle-level essays a 9 or 10. If you are already writing at the strong level, Claude Opus 5 will confirm it slightly more often than the engine. If you are not, it is the model most likely to tell you that you are.
The engine is the only grader in the test that is within ten points of the leader at every level, and the only one at 100% in the middle. That is the profile a candidate needs: a grader that is not brilliant at one level and wrong at another, but right nearly everywhere and right where most candidates are.
How the engine grades a response
Here is what happens in the minute or so between pressing submit and reading a score.
- The rule layer reads the essay first. Before any model sees it, the engine counts the words, checks the email for a greeting and a closing, checks the survey response for a clear choice, and matches the essay against the required points in the prompt. These checks become hard limits the rater cannot override: a survey response that never picks a side, or an email that reads like a survey answer, cannot score above the middle of the scale on task fulfillment.
- The essay is compared with real graded samples. The rater sees three real CELPIP responses to the same kind of task, one at the weak level, one in the middle, one at the top, all with their real slips in them, and places the essay relative to them on each of the four dimensions: content and coherence, vocabulary, readability, task fulfillment. Comparable to the middle sample is a 7 or 8; comparable to the top sample is a 10 or 11; clearly stronger than the top sample is a 12.
- Every judgement needs evidence. For each dimension the rater must quote the phrases from the essay that support the score before it gives the score. Quotes that cannot be found in the essay are discarded.
- Missed points cap the score. If the rater reports a required point as missing, the task-fulfillment score is held to 8 or below, so the number and the feedback agree.
- The four dimensions become one score. The overall score is the average of the four, after the rule-layer limits have been applied.
- Feedback is built from the same evidence. Corrections, a model answer at the right length, the missed points by name and a limiting factor per dimension, all quoted from the writer’s own text. The section below lists what changed.
We also tested moving the engine to the two-grader, vote-and-reconcile design that our speaking scorer uses. It did not help writing: the official writing levels sit closer together on the 0 to 12 scale than the speaking levels do, and forcing a rater to choose between samples with no in-between option pushed middle-level essays up to a 9. The writing engine keeps the design that scores the middle correctly.
What is new in your writing feedback
The score is only half of what you get back. The writing feedback layer was rebuilt alongside the scorer, and everything in it is anchored to your own text. Here is what changed.
- Corrections you can find in your own essay. Every grammar, spelling and vocabulary correction quotes the exact words from your response, so the app can highlight them where you wrote them. Anything the checker cannot find word for word in your essay is dropped before you see it.
- A model answer, built from your answer. You get a rewritten version of your own response that keeps your ideas, covers every required point and lands in the 150 to 200 words the task asks for. If the first rewrite misses that length, it is redone once before it is shown.
- The required point you missed, by name. For the email task, the feedback lists the bullet points from the prompt that your response did not cover. A missed point also holds the task-fulfillment score to 8 or below, so the number and the feedback never disagree.
- One limiting factor per dimension. For each of the four dimensions, the feedback names the single thing holding that score down, in concrete terms, or says plainly that nothing is and what is carrying it. Deciding a dimension has no problem is a real answer, not a gap.
- Criticism in the right place. A weak tone is a task-fulfillment problem, a vague word choice is a vocabulary problem, a run-on sentence is a readability problem. Each point is filed under the dimension that owns it instead of being lumped into the overall summary.
- Sentence-level rewrites. Specific sentences from your essay come back with an improved version and one line on why it is better, so you can see the change rather than read about it.
- Your strongest and weakest dimension, side by side. The feedback tells you which of the four dimensions is carrying your score and which is holding it back, so you know what to practise next without decoding four numbers.
Every quotation in the feedback is checked against your essay before it is displayed. If the checker cannot find the words, the line is removed. That is why the feedback never tells you to fix something you did not write.
What to do with the score
A score you can trust changes how you practise. With a chatbot, you are never sure whether a move from 8 to 9 is you or the model, so every score has to be checked against a feeling. With a stable, accurate grader, the score is the feedback.
- Write the same task twice. Submit, read the missed points and the limiting factor for each dimension, rewrite with the model answer beside you, and submit again. A change of a point or more between the two is a change in your writing, because the engine gives the same essay the same score.
- Track the weakest dimension, not the total. The overall score is an average of four. Two candidates with an 8 can have completely different problems. The feedback names the dimension holding you back, and that is the one to work on.
- Treat a 9 as a near miss. Both of the engine’s misses on strong essays were a 9. A 9 from the engine means one developed reason or one precise word choice short of a 10.
- Alternate the tasks. Most candidates practise the task they like. The engine scores email and survey responses against different samples and different rules, so a strong email score says nothing about your survey response.
Unlimited practice for CA$25 a month
Accuracy is only useful if you can afford to use it every day. A writing score you can trust is most valuable on your thirtieth essay, not your third, because that is when it starts telling you whether a change in how you write is working. So the engine is priced to be used without counting.
Every Prepamigo plan includes unlimited scoring, unlimited attempts, and the full question bank: 80 complete practice sets and more than 2,000 practice tasks across Writing, Speaking, Reading and Listening, written to follow the task types, timing and format of the CELPIP General test. You write, you submit, and about a minute later you have the score, the corrections and a model answer.
Prepamigo
CA$25/ month
CA$24.98 month to month · CA$8.25 a month on the annual plan (CA$98.98 a year) · tax included · cancel any time
- Unlimited scoring for writing and speaking, with no daily, session or weekly cap.
- 80 practice sets and 2,000+ tasks across all four sections, modelled on the real test format, so you never have to write your own prompts.
- Write and submit. Scoring, corrections, the missed-point check and the model answer are handled for you.
- Feedback in your own words, with every quote checked against what you actually wrote.
Claude Max, for comparison
US$100/ month
about CA$138 before tax · 20x tier US$200 a month · Pro US$20 with far tighter caps
- Even Max is capped. A rolling five-hour usage window and a weekly limit; when you hit either, you wait.
- No question bank. You bring your own tasks, decide for yourself whether they resemble the real test, and keep track of what you have already practised.
- No required-point check, no model answer. You get a number and a paragraph.
- A plain request, not a calibrated grader. Asked in plain words, Claude Opus 5 matched verified scores 61% of the time and Claude Sonnet 5 56%, against 86% for the engine.
Claude plan prices and usage terms are as published on claude.com in September 2026 and may change. CELPIP is a trademark of its owner; Prepamigo is an independent practice platform and is not affiliated with, endorsed by or connected to the test maker. Prepamigo practice tasks are original material written to follow the public test format.
Availability
The Prepamigo Writing Scoring Engine is live now for all Prepamigo users on both writing tasks. There is nothing to enable. Every email and survey response is scored by it, and every score comes with feedback drawn from the writer’s own text.
For the head-to-head comparisons, read Prepamigo vs Claude for writing and Prepamigo vs ChatGPT for writing. For the speaking side, read Scoring Engine 2.0 for Speaking. Or write a response and watch the engine run.
