Share of essays scored at the verified level
36 official CELPIP writing responses, 12 at each of three levels, across both writing tasks. Claude was given the task and the essay and asked, in plain words, to score it on the CELPIP scale. Prepamigo's writing scorer ran as it does in production.
86%
Prepamigo Writing Scorer
61%
Claude Opus 5
56%
Claude Sonnet 5
Prepamigo vs. the top chatbot plans
Canadian dollars. Prepamigo prices include tax. Claude Max (from US$100) and ChatGPT Pro (US$200) converted at 1.38, before tax. Accuracy is the share of 36 official essays scored at the verified level when the named model is asked, in plain words, to score the essay; single run.
Writing is the skill where a chatbot looks most convincing as a grader. There is no recording to transcribe, no accent to worry about, just text going in and a number coming out, with a paragraph of confident explanation. So we tested the number. We gave Claude the same 36 official CELPIP writing responses our own scorer is measured on, twelve at each level, and asked it to score them the way a student would ask.
The short answer: Prepamigo’s writing scorer landed on the verified level 86% of the time. Claude Opus 5 landed there 61% of the time and Claude Sonnet 5 56% of the time. The gap is not spread evenly. On weak essays and strong essays Claude is respectable, and at the top Opus 5 is actually better than we are. On middle-level essays, the 7s and 8s that most candidates actually write, Opus 5 was right one time in four and Sonnet 5 one time in three. Prepamigo was right on all twelve.
The rest of this article walks through those results level by level, explains why a general-purpose model loses the middle of the scale, looks at what the feedback on each side is worth, and compares what it costs to practise writing seriously with each. We have tried to be fair to Claude throughout, including the one place it beat us.
The plain test: what a student actually gets
Eighty-six percent against sixty-one and fifty-six. On 36 essays, that is nine more correct scores than Opus 5 and eleven more than Sonnet 5. Put another way, Prepamigo makes 64% fewer scoring mistakes than Opus 5 and 69% fewer than Sonnet 5 on the same essays.
What Claude was given matters, so here it is. Each model was told it was an experienced CELPIP writing examiner, given the task exactly as the candidate saw it, including the required points for the email task and the options for the survey task, and given the essay. It was asked for one score from 0 to 12. That is the whole prompt. It is what you would type into a chat window, and it is a fairer test than a one-line “grade this” because the model at least knows what the task asked for.
Prepamigo’s scorer received the same task and the same essay. It also had what it always has in production: real graded sample essays for that task at each level to compare against, and a rule layer that checks the things a rater checks first, such as whether every required point is there, whether a survey response actually picks a side, whether an email has a greeting and a closing, and how long the response is.
Where the gap opens: level by level
An overall figure hides where the errors fall, and in writing the errors fall in one place. Here are the same results split by verified level: weak essays scored 4 or 5, middle essays scored 7 or 8, and strong essays scored 10 and above.
Weak essays (verified score 4 or 5)
12 official essays. Claude's misses here are mostly a 3, one level below.
75%
Prepamigo Writing Scorer
67%
Claude Opus 5
50%
Claude Sonnet 5
On weak essays, Prepamigo is at 75%, Opus 5 at 67%, Sonnet 5 at 50%. Prepamigo’s three misses all went one notch too high, to a 6 or 7, on essays that were short but tidy. Claude’s misses go the other way: Sonnet 5 scored four of the twelve weak essays a 3, and Opus 5 scored two of them a 3. A 3 on this scale is a response that barely engages with the task. These essays engage with it; they are just thin. Claude, with no example of what a 4 looks like, reads thin as failing.
Middle essays (verified score 7 or 8)
12 official essays. This is the level most candidates are at, and the level where Claude fails.
100%
Prepamigo Writing Scorer
33%
Claude Sonnet 5
25%
Claude Opus 5
The middle is the whole story. Prepamigo scored all twelve middle-level essays inside the 7 to 8 band. Opus 5 got three of them; Sonnet 5 got four. Opus 5’s misses were almost all upward: it gave a 9 to six of the twelve and a 10 to one. Sonnet 5 split both ways, with four essays pushed down to a 6 and four pushed up to a 9 or 10.
Think about what that means for the candidate who wrote one of those essays. A 7 or 8 is the most common real score on the test, and it is the score that decides whether you have reached your immigration target or not. Ask Opus 5, and more than half the time it will tell you that you are a 9 or a 10, ready to stop practising. You are not. Ask Sonnet 5, and one time in three it will tell you that you have slipped to a 6. You have not. Neither answer helps you decide what to do next.
Strong essays (verified score 10 or above)
12 official essays. The one level where Claude Opus 5 beats Prepamigo.
92%
Claude Opus 5
83%
Prepamigo Writing Scorer
83%
Claude Sonnet 5
At the top, Opus 5 is the best grader in this test: eleven of twelve strong essays recognised. Prepamigo and Sonnet 5 each recognised ten. Prepamigo’s two misses were both a 9, one notch under; Sonnet 5’s were a 9 and a 7. We report this because it is true, and because it explains the overall pattern: Claude is generous. It hands out 9s, 10s and 11s freely, which happens to be right on strong essays and wrong on everything else.
Why a chatbot loses the middle of the scale
A middle-level CELPIP essay is a specific thing. It covers the task, it is organised, it has few errors that block understanding, and it is also generic: the reasons are asserted rather than developed, the vocabulary is safe, the sentences are all the same shape. A rater who has seen hundreds of these knows exactly where it sits. A general-purpose model has not seen hundreds of these labelled as 7s. It has seen a clean, organised, mostly correct piece of writing, and it reaches for a high number.
That is the whole difference. Prepamigo’s scorer does not judge an essay against an idea of good writing. It judges it against real graded essays for the same task: one at the weak level, one in the middle, one at the top, all of them real responses with real slips in them. A clean but generic essay lands next to the middle sample because that is what it most resembles. A thin essay lands next to the weak sample rather than below it, because the weak sample is also thin and still a 4.
On top of the comparison sits a rule layer that handles the things a model is bad at counting. Did the email cover all three required points, or only two? Did the survey response actually choose an option, or hedge between them? Is there a greeting and a closing? Is the response 180 words or 60? A missed required point holds the task-fulfillment score to 8 or below no matter how good the English is, because that is how the test works. Claude, asked plainly, has no such floor or ceiling; it decides everything by impression.
We tested whether telling Claude about the levels would fix this. It does not, on its own. Giving a model a rubric produces the same drift. What moves the number is giving it the actual graded samples and a rule layer, which is another way of saying: building a scorer.
The workflow gap: what you have to bring
Writing is the skill where Claude is easiest to use as a grader, and it is worth being honest about that. You paste the essay, you get a score in seconds, and if you want it to explain itself it will, at length. There is no transcription step, no audio, no setup. For a quick second opinion it is genuinely convenient.
What you do not get is the task. Claude has no bank of CELPIP-style email and survey prompts with the right required points, options and word counts, so you either write your own prompt and hope it resembles the test, or you practise on whatever you can find. You do not get a required-point check, because Claude does not know which points the real test would list. You do not get a model answer rewritten from your own essay at the right length. And you do not get a number that was calibrated against real graded essays, which is why the middle of the scale goes wrong.
On Prepamigo, you open a task from the bank, the timer runs, you write, and about a minute after you submit you have four dimension scores, an overall score, a list of anything you missed, corrections quoted from your own text, and a rewritten model answer. The fortieth essay of the month costs you the same effort as the first.
The same essay, the same score
A score you cannot reproduce is not a measurement. If the same essay gets an 8 today and a 10 tomorrow, you have learned nothing about your writing. So after the headline run we scored all 36 essays twice more on different days. Prepamigo’s writing scorer scored 86%, 89% and 86% across the three runs, with the same twelve middle-level essays inside the band every time.
Claude’s runs moved too, but the shape of its errors did not: on every run, most of the middle-level essays came back as a 9 from Opus 5, and the weak essays kept collecting 3s from Sonnet 5. That consistency is the useful finding. The problem is not that Claude is noisy. It is that Claude is confidently miscalibrated in the same direction every time, which is harder to notice.
Feedback: a good explanation of the wrong number
Claude writes excellent feedback prose. It is clear, patient and specific, and if you ask it to explain why an essay is a 9 it will produce a convincing paragraph. The trouble is that the paragraph is written to justify a number it has already chosen, and when the number is wrong the explanation is wrong with it. An essay that is really a 7 gets praised for qualities it does not have, and the candidate walks away reassured.
Prepamigo’s feedback works from the other end. Every correction quotes the exact words from your essay, and anything the checker cannot find in your text is dropped before you see it. A missed required point is named. Each dimension gets one limiting factor, or a plain statement that nothing is holding it back. And you get a model answer rewritten from your own response at the length the task asks for, so you can see the gap rather than read about it. The section below lists what changed in the feedback layer.
What is new in your writing feedback
The score is only half of what you get back. The writing feedback layer was rebuilt alongside the scorer, and everything in it is anchored to your own text. Here is what changed.
- Corrections you can find in your own essay. Every grammar, spelling and vocabulary correction quotes the exact words from your response, so the app can highlight them where you wrote them. Anything the checker cannot find word for word in your essay is dropped before you see it.
- A model answer, built from your answer. You get a rewritten version of your own response that keeps your ideas, covers every required point and lands in the 150 to 200 words the task asks for. If the first rewrite misses that length, it is redone once before it is shown.
- The required point you missed, by name. For the email task, the feedback lists the bullet points from the prompt that your response did not cover. A missed point also holds the task-fulfillment score to 8 or below, so the number and the feedback never disagree.
- One limiting factor per dimension. For each of the four dimensions, the feedback names the single thing holding that score down, in concrete terms, or says plainly that nothing is and what is carrying it. Deciding a dimension has no problem is a real answer, not a gap.
- Criticism in the right place. A weak tone is a task-fulfillment problem, a vague word choice is a vocabulary problem, a run-on sentence is a readability problem. Each point is filed under the dimension that owns it instead of being lumped into the overall summary.
- Sentence-level rewrites. Specific sentences from your essay come back with an improved version and one line on why it is better, so you can see the change rather than read about it.
- Your strongest and weakest dimension, side by side. The feedback tells you which of the four dimensions is carrying your score and which is holding it back, so you know what to practise next without decoding four numbers.
Every quotation in the feedback is checked against your essay before it is displayed. If the checker cannot find the words, the line is removed. That is why the feedback never tells you to fix something you did not write.
What it costs to practise every day
A writing score is most useful on your thirtieth essay, not your third. That is when it starts telling you whether a change in how you write is working. So the practical question is what each option costs to use in volume.
Claude Pro is US$20 a month, about CA$28, or US$17 a month on annual billing, as published on claude.com in September 2026, with a rolling five-hour usage window and a weekly limit. The Max plans, from US$100 a month, about CA$138 before tax, raise those limits but do not remove them. None of the plans include a task bank, a timer, a required-point check, a model answer, or a score calibrated against real graded essays.
Prepamigo is CA$24.98 a month, or CA$8.25 a month on the annual plan, which is CA$98.98 for the year, tax included, and you can cancel at any time. Every plan includes unlimited scoring with no daily, session or weekly cap, unlimited attempts, and the full question bank: 80 complete practice sets and more than 2,000 practice tasks across Writing, Speaking, Reading and Listening, written to follow the task types, timing and format of the CELPIP General test.
When Claude is the better tool
- Polishing a strong essay. If you are already at a 10 and want to know what an 11 or 12 would need, Opus 5 recognises strong writing reliably and explains well. That is the one level where it beat us.
- Rewriting a sentence five ways. Asking for alternatives to a clumsy sentence is quick and the suggestions are good.
- Talking it through. If you want to argue about why a reason is weak, a conversation is the right shape. Prepamigo gives you a verdict and evidence; it does not chat.
- Anything outside the test. Claude is a general assistant and Prepamigo is not.
What Claude should not be, on this evidence, is the thing that tells a 7 or 8 writer where they stand. That is most candidates, and it is exactly where Claude is wrong three times in four.
FAQs
Can Claude score my CELPIP writing? It will give you a number. On 36 official responses with verified scores, Claude Opus 5 was right 61% of the time and Claude Sonnet 5 56%, against 86% for Prepamigo. On middle-level essays: 25% and 33% against 100%.
Is Prepamigo more accurate than Claude for writing? Overall, yes, by 25 to 30 points. On strong essays Opus 5 is better, 92% against 83%. On middle-level essays, where most candidates are, the gap is 100% against 25%.
Why does Claude give my average essay a 9? With nothing to compare against, a clean, organised, generic essay reads as strong. Opus 5 gave a 9 or 10 to seven of twelve verified 7-to-8 essays. Prepamigo compares every essay with real graded samples for the same task.
How much does each cost? Claude Max starts at US$100 a month, about CA$138 before tax; Pro is US$20 with tighter caps. Prepamigo is CA$24.98 a month tax included, or CA$8.25 a month on the annual plan, with unlimited scoring and 80 practice sets.
For the same comparison against ChatGPT, read Prepamigo vs ChatGPT for writing. For how the writing scorer works, read inside the Prepamigo writing scorer. Or write a response and see the score.
