Prepamigo writing scorer on 36 official essays
Share of essays scored inside the verified level, overall and by level. Twelve essays per level across both writing tasks; the six sample essays the scorer uses for calibration are excluded.
86%
Overall
75%
Weak (4–5)
100%
Middle (7–8)
83%
Strong (10+)
Prepamigo vs. the top chatbot plans
Canadian dollars. Prepamigo prices include tax. Claude Max (from US$100) and ChatGPT Pro (US$200) converted at 1.38, before tax. Accuracy is the share of 36 official essays scored at the verified level when the named model is asked, in plain words, to score the essay; single run.
It is the right question to ask of any practice tool, and most tools do not answer it. A writing score is only worth something if it tells you what a real rater would say, and the only way to know whether it does is to measure. So this article is the measurement, laid out plainly, with the parts that flatter us and the parts that do not.
The answer in one paragraph: on 36 official CELPIP writing responses, each with an independently verified score, Prepamigo’s writing scorer landed on the verified level 86% of the time. It got every middle-level essay right, 75% of weak essays and 83% of strong ones. It produced the same result, give or take one essay, on three separate runs. Asked in plain words to score the same essays, the best frontier model, GPT 5.6 sol, reached 78%; the Claude models reached 61% and 56%.
What follows is what the numbers mean at each level, where the errors fall and in which direction, how stable the score is, how it compares with the chatbots you might be using instead, what is new in the feedback, and how to read your own score in light of all of it.
What “accurate” means here
We are not claiming that a Prepamigo score predicts your test result. No practice tool can promise that. What we measure is narrower and more honest: given an essay whose score has been independently verified, how often does the scorer produce a score at the same level?
CELPIP writing is reported on a scale that runs to 12, and the verified scores in our reference data come in three bands: weak, meaning 4 or 5; middle, meaning 7 or 8; and strong, meaning 10 and above. We count the scorer as correct when its score lands inside the verified band. A strong essay scored a 10, 11 or 12 is correct; scored a 9 it is a miss, even though it is a near miss.
Accuracy at each level
Weak essays: 75%. Nine of twelve inside the band. All three misses went one notch too high: two essays scored a 6 and one a 7. Each was short and thin but tidy, with few errors, and tidiness bought it a point it had not earned on content. This is the scorer’s characteristic error at the low end, and it is a generous one: a weak writer is told they are a little better than they are, never worse.
Middle essays: 100%. Twelve of twelve inside the 7 to 8 band, on all three runs. This is the level most candidates are at and the level that decides most immigration targets, and it is where the scorer is strongest. It is also where every chatbot we tested is weakest, for reasons the comparison section explains.
Strong essays: 83%. Ten of twelve. Both misses were a 9, one notch under. A strong CELPIP essay is not flawless, and the scorer occasionally reads one slip too many. Neither miss was a 7 or 8; a strong essay is never told it is average.
By task, the scorer was right on 89% of email responses and 83% of survey responses. Survey responses are harder to place because the task rewards a clear position and developed reasons, and a response that hedges between the options is penalised by rule.
Which way the errors go
A grader that is wrong 14% of the time can be wrong in a harmless way or a harmful way. Harmful is telling a candidate who needs an 8 that they have it when they do not. Harmless is telling them they are a point short when they are not; that costs a week of practice and nothing else.
Of the scorer’s five misses on 36 essays, three were weak essays scored a 6 or 7 and two were strong essays scored a 9. No middle-level essay was scored high, and no weak essay was scored strong. In practical terms: if the scorer says you have reached the middle level, you have, or you are one point short. If it says you have reached the strong level, you have. The only place it flatters is at the bottom, where a short, tidy essay can pick up a point it did not earn.
Compare that with the chatbots on the same essays. Claude Opus 5 called seven middle-level essays a 9 or 10, which is the harmful direction for the candidates it affects most. GPT 5.6 sol called three middle-level essays a 6, which is the harmless direction, and held three strong essays to a 9. A grader’s accuracy number tells you how often it is wrong; the direction tells you what it costs you when it is.
Email versus survey
The 36 essays were split between the two writing tasks, and the scorer did slightly better on the email task, at 89%, than on the survey response, at 83%. That is one extra miss on the survey side, and it is in the direction you would expect.
Email responses are scored largely on coverage: were the required points addressed, is the tone right for the reader, is there a greeting and a closing. Those are things the scorer checks by rule before any judgement is made, so an email that has them is placed reliably. Survey responses are scored on the quality of an argument: a clear position and reasons that are developed rather than listed. Development is a judgement, and judgement is where a grader, human or otherwise, has room to differ by a point.
For you, this means a survey score is very slightly softer than an email score. If you are hovering at the edge of your target on survey responses, write two rather than one before you decide where you stand.
From your side of the screen
The level view tells you how the scorer does on an essay of a known level. You are looking at it the other way round: you have a score and want to know how much to believe it.
- If it says 4 or 5: the essay was at the weak level nine times in nine. A low score from the scorer is right.
- If it says 7 or 8: the essay was at the middle level twelve times in fifteen; the other three were weak essays scored a notch high. A 7 or 8 means middle, and possibly a little lower.
- If it says 10 or above: the essay was at the strong level ten times in ten. A 10 from the scorer is a 10.
- If it says 9: both 9s in the test were verified strong essays. A 9 is the score to read carefully: it means close to the top, and one developed reason or one precise word choice short of it.
Does the same essay get the same score?
Accuracy without consistency is luck. So we scored all 36 essays three times on different days with nothing changed in between. The scorer returned 86%, 89% and 86%. The same twelve middle-level essays were inside the band every time, and no essay moved by more than one point between runs.
The consistency comes from how the scoring model is run: reasoning switched off, a fixed temperature, and a comparison against fixed graded samples rather than a free judgement. For you, it means a change in your score is a change in your writing. If you rewrite an essay and it moves from 8 to 10, that is not the grader having a good day.
How this compares with what you might use instead
An accuracy number means more with something to compare it to. So we gave the same 36 essays to the chatbots people actually use to check their writing, the way a person would: the task, the essay, and a plain request for a score from 0 to 12.
Share of essays scored inside the verified level
Same 36 official essays. Each model was asked, in plain words, to score the essay; Prepamigo ran in its normal production configuration.
86%
Prepamigo Writing Scorer
78%
GPT 5.6 sol
69%
GPT 5.6 luna
61%
Gemini
61%
Claude Opus 5
58%
GPT-4o mini
56%
Claude Sonnet 5
The scorer leads the best chatbot by eight points and the Claude models by 25 to 30. The shape of the errors is what matters. On middle-level essays, Prepamigo is at 100%, GPT 5.6 sol at 75%, Gemini 42%, Claude Sonnet 5 33% and Claude Opus 5 25%: a chatbot with nothing to compare against reads a clean, generic 7 as a 9, or a generic 7 with a few errors as a 6. On strong essays Claude Opus 5 is actually the best grader in the test at 92% against Prepamigo’s 83%; it is generous, which is right at the top and wrong everywhere else. GPT-4o mini gave a 9 to nine of twelve strong essays.
What the scorer has that the chatbots do not is real graded essays for the same task at each level to compare against, and a rule layer that checks required points, survey stance, greeting and closing, and length before any model gives an opinion. That is the difference between 86% and 78%, and it is nearly all in the middle of the scale.
What is new in your writing feedback
The score is only half of what you get back. The writing feedback layer was rebuilt alongside the scorer, and everything in it is anchored to your own text. Here is what changed.
- Corrections you can find in your own essay. Every grammar, spelling and vocabulary correction quotes the exact words from your response, so the app can highlight them where you wrote them. Anything the checker cannot find word for word in your essay is dropped before you see it.
- A model answer, built from your answer. You get a rewritten version of your own response that keeps your ideas, covers every required point and lands in the 150 to 200 words the task asks for. If the first rewrite misses that length, it is redone once before it is shown.
- The required point you missed, by name. For the email task, the feedback lists the bullet points from the prompt that your response did not cover. A missed point also holds the task-fulfillment score to 8 or below, so the number and the feedback never disagree.
- One limiting factor per dimension. For each of the four dimensions, the feedback names the single thing holding that score down, in concrete terms, or says plainly that nothing is and what is carrying it. Deciding a dimension has no problem is a real answer, not a gap.
- Criticism in the right place. A weak tone is a task-fulfillment problem, a vague word choice is a vocabulary problem, a run-on sentence is a readability problem. Each point is filed under the dimension that owns it instead of being lumped into the overall summary.
- Sentence-level rewrites. Specific sentences from your essay come back with an improved version and one line on why it is better, so you can see the change rather than read about it.
- Your strongest and weakest dimension, side by side. The feedback tells you which of the four dimensions is carrying your score and which is holding it back, so you know what to practise next without decoding four numbers.
Every quotation in the feedback is checked against your essay before it is displayed. If the checker cannot find the words, the line is removed. That is why the feedback never tells you to fix something you did not write.
How to read your score
- A 4 or 5 is right. Read the corrections; they are quoted from your own text. Then read the missed points and the model answer, and write the same task again.
- A 7 or 8 is middle, possibly a little lower. Look at the limiting factor for each dimension. The one that names a concrete problem is the one to work on.
- A 9 is close. Compare your essay with the model answer paragraph by paragraph. The gap is usually one undeveloped reason or one generic word choice.
- A 10 or above is a 10. You are ready on that task type. Move to the one you avoid.
- Trust changes more than levels. Because the score is stable, a change across attempts at the same task is a change in your writing.
A practice routine that uses the number
The point of an accurate, stable score is that you can stop second-guessing it and start using it. Here is the routine we suggest, built around what the numbers on this page say the score can and cannot tell you.
- Write one email and one survey response, cold. Do not look at the model answer first. Submit both. The two scores together are your starting level; if they differ by more than a point, the lower one is the task to practise.
- Read the missed points and the limiting factor before the corrections.Corrections fix sentences; missed points and limiting factors fix scores. A missing required point is worth more than every comma.
- Rewrite the same response with the model answer open. Keep your ideas, borrow its structure. Submit again. Because the scorer gives the same essay the same score, the difference between the two submissions is entirely your rewrite.
- Move to a new task when the rewrite holds. A single 10 on a rewrite is not yet a level. Two 10s on two different tasks are.
- Read a 9 as a near miss and a 6 as a real 6. The scorer does not pull middle essays down. If it says 6, the essay is thin, and the limiting factor will say where.
FAQs
How accurate is Prepamigo’s writing score? 86% at the exact level on 36 official essays: 75% weak, 100% middle, 83% strong. The same result, give or take one essay, on three runs.
Is it the same as my real score? No practice tool can promise that. What we measure is how often the scorer agrees with a verified score on the same essay.
Why did I get a 9 on a strong essay? That is the scorer’s most common miss at the top: two of twelve strong essays scored a 9. Rewrite with the model answer beside you and score again.
Is it better than ChatGPT or Claude? 86% against 78% for GPT 5.6 sol and 61% and 56% for the Claude models. On middle-level essays, 100% against 75% for the best of them.
To see how the scorer works, read inside the Prepamigo writing scorer. To see the same test on all six chatbots, read which AI scores CELPIP writing most accurately. Or write a response and see the numbers for yourself.
