Share of essays scored at the verified level
36 official CELPIP writing responses, 12 per level, both writing tasks. Each model was given the task and the essay and asked, in plain words, to score it on the CELPIP scale. Prepamigo's writing scorer ran in its normal production configuration.
86%
Prepamigo Writing Scorer
78%
GPT 5.6 sol
69%
GPT 5.6 luna
61%
Gemini
61%
Claude Opus 5
58%
GPT-4o mini
56%
Claude Sonnet 5
Prepamigo vs. the top chatbot plans
Canadian dollars. Prepamigo prices include tax. Claude Max (from US$100) and ChatGPT Pro (US$200) converted at 1.38, before tax. Accuracy is the share of 36 official essays scored at the verified level when the named model is asked, in plain words, to score the essay; single run.
Plenty of CELPIP candidates check their writing with an AI chatbot. Paste the email, ask for a score, read the paragraph. For writing it is easy: no recording, no transcript. What nobody tells you is how often the number is right, or which model to trust, or which mistakes each one makes. We wanted to know, so we gave six chatbots and our own scorer the same 36 official essays with verified scores and counted.
The one-sentence version: GPT 5.6 sol is the most accurate chatbot for CELPIP writing at 78%, the Claude models are the least accurate at 61% and 56% because they inflate middle-level essays, and a purpose-built scorer, Prepamigo’s, reached 86% on the same essays and got every middle-level essay right. We should be upfront that Prepamigo is our product. Where a chatbot beat us on a particular level, we say so; two of them did.
The leaderboard
Three tiers are visible in the chart above. Prepamigo at 86% stands alone. GPT 5.6 sol at 78% and GPT 5.6 luna at 69% form a second tier: usable, with known habits. Everything else is at or near 60%, which for a three-band scale is not much better than an informed guess.
The pattern is different from speaking, where the Claude models did well and the GPT models did badly. In writing it is reversed. The GPT models are slightly severe: they pull a generic 7 down to a 6 and hold a strong essay with a slip to a 9. The Claude models are generous: they push a generic 7 up to a 9 or 10 and score thin weak essays a 3. Gemini behaves like Claude in the middle. GPT-4o mini treats 9 as the ceiling.
Results at each level
Weak essays (verified 4 or 5)
12 official essays per system.
83%
GPT 5.6 sol
83%
GPT 5.6 luna
83%
GPT-4o mini
75%
Prepamigo Writing Scorer
67%
Claude Opus 5
58%
Gemini
50%
Claude Sonnet 5
Weak essays are the easy end of the scale, and the three GPT models tie at 83%, ahead of Prepamigo at 75%. Prepamigo’s three misses all went one notch high, to a 6 or 7, on short but tidy essays. The Claude models miss in the other direction: Sonnet 5 scored four of twelve weak essays a 3, a level below, and Opus 5 scored two of them a 3. A thin essay that still does the task is a 4, not a 3, and a model with no example of a 4 cannot tell the difference.
Middle essays (verified 7 or 8)
12 official essays per system. The level most candidates are at.
100%
Prepamigo Writing Scorer
75%
GPT 5.6 sol
67%
GPT-4o mini
58%
GPT 5.6 luna
42%
Gemini
33%
Claude Sonnet 5
25%
Claude Opus 5
The middle band separates the field. Prepamigo scored all twelve. GPT 5.6 sol scored nine, pulling three down to a 6. GPT-4o mini scored eight and pushed four up to a 9. Luna scored seven, mostly pulling down. Gemini scored five and pushed six up to a 9 or 10. Claude Sonnet 5 scored four, split between 6s and 9s. Claude Opus 5 scored three and gave a 9 or 10 to seven of the twelve.
If you are a 7 or 8 writer, and most candidates are, this is the chart to remember. Ask Claude Opus 5 and more than half the time it will tell you that you are done. Ask Gemini and half the time it will say the same. Ask GPT 5.6 sol and one time in four it will tell you that you have slipped to a 6.
Strong essays (verified 10 or above)
12 official essays per system. Nearly every miss is a 9.
92%
Claude Opus 5
83%
Prepamigo Writing Scorer
83%
Gemini
83%
Claude Sonnet 5
75%
GPT 5.6 sol
67%
GPT 5.6 luna
25%
GPT-4o mini
At the top, Claude Opus 5 is the best grader in the test with eleven of twelve strong essays recognised; it gave five of them an 11. Prepamigo, Gemini and Sonnet 5 recognised ten each. GPT 5.6 sol recognised nine and luna eight, holding the rest to a 9. GPT-4o mini recognised three and gave a 9 to nine: it does not hand out 10s. Claude’s generosity is right here and wrong everywhere else; the GPT models’ severity is wrong here and roughly right elsewhere.
Why the models disagree about the middle
A middle-level CELPIP essay is a specific thing: it covers the task, it is organised, it has few errors that block understanding, and it is generic, with reasons asserted rather than developed and vocabulary that is safe rather than precise. A trained rater has seen hundreds of these and knows where they sit. A general-purpose model has seen a tidy, organised, mostly correct piece of writing and has to decide from impression. Claude’s impression is that tidy means strong. GPT’s impression is that generic means weak. Both are wrong about a 7, in opposite directions.
Prepamigo’s scorer does not judge from impression. It compares the essay with real graded responses to the same kind of task at each level, all of them real essays with real slips, and places it next to the one it most resembles on each of the four dimensions. A clean but generic essay lands next to the middle sample, because that is what it is. On top of the comparison sits a rule layer for the things a model is bad at counting: whether every required point is covered, whether a survey response picks a side, whether the email has a greeting and a closing, and how long it is. A missed required point holds task fulfillment to 8 or below regardless of the English, because that is how the test works.
We also tried giving the writing scorer the two-grader, vote-and-reconcile design our speaking scorer uses. It made writing worse, because the official writing levels sit only two points apart on the 0 to 12 scale and a forced choice between samples pushed middle essays up to a 9. Writing keeps the design that gets the middle right.
Why the ranking is the reverse of speaking
If you have read our speaking comparison, the writing leaderboard will look upside down. On speaking transcripts Claude Opus 5 was the best chatbot at 62% and GPT 5.6 sol the worst of the large models at 35%. On writing it is GPT 5.6 sol at 78% and the Claude models at the bottom.
The reason is what each model is generous about. A CELPIP speaking transcript is spoken English written down: fragments, repetitions, self-corrections. GPT reads that as broken and scores it harshly, which on a transcript is wrong; Claude reads through it to the ideas. A CELPIP essay is the opposite: it is edited, organised, tidy on the surface, and often generic underneath. Claude reads the tidiness as strength and inflates the middle; GPT reads the generic content as weakness and is roughly right, or a point low.
The practical lesson is that there is no single best chatbot for CELPIP. The model that grades your speaking well is the one that will flatter your writing, and the other way round. A purpose-built scorer sidesteps the question by comparing against graded samples of the right kind for each task. Prepamigo’s speaking engine and writing engine are separate systems with different designs, because the two tasks fail differently.
The paragraph that comes with the number
Every chatbot returns a paragraph of feedback with its score, and the paragraph is usually more confident than the number deserves. Three things to check before you act on it.
- Does it quote you? A correction that does not point to your exact words is a general rule, not feedback on your essay. Most chatbot advice is of the kind that applies to any essay: vary your sentence structure, use more precise vocabulary, develop your reasons. True, and not actionable.
- Does it name the missing point? If you pasted the task, a good response tells you which required point you did not cover. If you did not paste it, the chatbot cannot know, and it will praise the coverage anyway.
- Does the criticism match the score? A chatbot paragraph often lists three or four problems and then gives a 10, or praises the essay and gives a 6. When the words and the number disagree, neither is reliable.
Prepamigo’s feedback is built the other way round: the scorer must quote the phrases that support each dimension score before it gives the score, and a quote that cannot be found in your essay is discarded. The missed points are listed by name, and a missed point caps the task-fulfillment score, so the words and the number cannot disagree. The section below lists what is in it.
Notes on each model
- GPT 5.6 sol. The best chatbot for writing at 78%, and the one to use if you use one. Slightly severe: three middle essays pulled to a 6, three strong essays held to a 9. On weak essays it is better than Prepamigo, 83% against 75%.
- GPT 5.6 luna. 69%. Same shape as sol but more severe in the middle, where it scored seven of twelve. Good on weak essays.
- Gemini. 61%. Fine at the extremes, poor in the middle at 42%, where it pushed six of twelve essays up to a 9 or 10.
- Claude Opus 5. 61% overall but the best grader in the test on strong essays at 92%. In the middle it is the worst at 25%, giving seven of twelve a 9 or 10. If your writing is already strong, Opus 5 will confirm it; if it is average, Opus 5 will tell you it is strong.
- GPT-4o mini. 58%. Treats 9 as the top of the scale: nine of twelve strong essays scored a 9. Do not use it to grade writing if you are anywhere near a 10.
- Claude Sonnet 5. 56%, the least accurate. Scored four weak essays a 3 and split the middle band between 6s and 9s.
If you are going to use a chatbot anyway
- Use GPT 5.6 sol. Not a Claude model for anything below a strong essay, and not a small model for anything above a weak one.
- Paste the whole task. The required points for the email task and the options for the survey task change the score. A model that does not know the points cannot tell you one is missing.
- Ask it to check the points first. Ask for a yes or no on each required point, then the score. This is the rule layer, done by hand.
- Ask twice, keep the lower answer. GPT 5.6 sol scored 78%, 83% and 81% on three runs of the same essays; a single answer carries a few points of luck.
- Read a 9 and a 6 with suspicion. From a GPT model, a 9 is often a held-back 10 and a 6 is often a pulled-down 7. From a Claude model, a 9 is often a promoted 7.
The one input you cannot supply yourself is a verified graded essay for the same task at each level, which is what a purpose-built scorer compares against. That, and the required-point check, is the difference between 78% and 86%.
What is new in your writing feedback
The score is only half of what you get back. The writing feedback layer was rebuilt alongside the scorer, and everything in it is anchored to your own text. Here is what changed.
- Corrections you can find in your own essay. Every grammar, spelling and vocabulary correction quotes the exact words from your response, so the app can highlight them where you wrote them. Anything the checker cannot find word for word in your essay is dropped before you see it.
- A model answer, built from your answer. You get a rewritten version of your own response that keeps your ideas, covers every required point and lands in the 150 to 200 words the task asks for. If the first rewrite misses that length, it is redone once before it is shown.
- The required point you missed, by name. For the email task, the feedback lists the bullet points from the prompt that your response did not cover. A missed point also holds the task-fulfillment score to 8 or below, so the number and the feedback never disagree.
- One limiting factor per dimension. For each of the four dimensions, the feedback names the single thing holding that score down, in concrete terms, or says plainly that nothing is and what is carrying it. Deciding a dimension has no problem is a real answer, not a gap.
- Criticism in the right place. A weak tone is a task-fulfillment problem, a vague word choice is a vocabulary problem, a run-on sentence is a readability problem. Each point is filed under the dimension that owns it instead of being lumped into the overall summary.
- Sentence-level rewrites. Specific sentences from your essay come back with an improved version and one line on why it is better, so you can see the change rather than read about it.
- Your strongest and weakest dimension, side by side. The feedback tells you which of the four dimensions is carrying your score and which is holding it back, so you know what to practise next without decoding four numbers.
Every quotation in the feedback is checked against your essay before it is displayed. If the checker cannot find the words, the line is removed. That is why the feedback never tells you to fix something you did not write.
FAQs
Which AI is most accurate at scoring CELPIP writing? Among chatbots, GPT 5.6 sol at 78%, then luna 69%, Gemini and Claude Opus 5 61%, GPT-4o mini 58%, Claude Sonnet 5 56%. Prepamigo’s writing scorer reached 86% on the same essays.
ChatGPT or Claude for writing? ChatGPT, clearly: 78% against 61% and 56%. The Claude models inflate middle-level essays to a 9 or 10. Opus 5 is the best of all on strong essays.
Why do they give my average essay a 9? Tidy and organised reads as strong to a model with nothing to compare against. Prepamigo compares with real graded samples for the same task.
How do I get a better score from a chatbot? Use GPT 5.6 sol, paste the full task, ask it to check each required point first, ask twice, and read 9s and 6s with suspicion.
For the two head-to-heads, read Prepamigo vs Claude for writing and Prepamigo vs ChatGPT for writing. For how the scorer works, read inside the Prepamigo writing scorer. Or write a response and skip the guesswork.
