Product

Prepamigo vs ChatGPT for CELPIP Writing: Accuracy, Limits and Cost

Sep 8, 2026

Share of essays scored at the verified level

36 official CELPIP writing responses, 12 at each of three levels, across both writing tasks. Each GPT model was given the task and the essay and asked, in plain words, to score it on the CELPIP scale. Prepamigo's writing scorer ran as it does in production.

86%

Prepamigo Writing Scorer

78%

GPT 5.6 sol

69%

GPT 5.6 luna

58%

GPT-4o mini

Prepamigo vs. the top chatbot plans

Canadian dollars. Prepamigo prices include tax. Claude Max (from US$100) and ChatGPT Pro (US$200) converted at 1.38, before tax. Accuracy is the share of 36 official essays scored at the verified level when the named model is asked, in plain words, to score the essay; single run.

Prepamigo
ChatGPT ProGPT 5.6 sol
Monthly price
CA$24.98 (inc. tax)
CA$276 (plus tax)
Scoring
Unlimited
Capped
Ready-made question bank
Yes
No
Practice sets
80
None
Practice tasks
2,000+
None
CELPIP format
Yes
No
Accuracy
86%
78%
Middle-level essays
100%
75%

ChatGPT is probably the most common tool CELPIP candidates use to check their writing. Paste the email, ask for a score, get a number and a paragraph. It is fast, it is on your phone, and for writing there is no transcription step in the way. So we tested the number. We gave three GPT models the same 36 official CELPIP writing responses our own scorer is measured on, twelve at each level, and asked them to score the essays the way a student would ask.

The short answer: Prepamigo’s writing scorer landed on the verified level 86% of the time. GPT 5.6 sol, the model behind ChatGPT’s paid plans, landed there 78% of the time. GPT 5.6 luna managed 69% and GPT-4o mini 58%. Eight points is a real gap, and it comes from two places: the middle of the scale, where Prepamigo scored all twelve essays correctly and GPT 5.6 sol scored nine, and the top, where GPT 5.6 sol held back a quarter of the strong essays to a 9.

We should say up front that GPT 5.6 sol is the best general-purpose writing grader we have tested, well ahead of any Claude model on the same essays. The rest of this article walks through the results level by level, explains where the eight points come from, looks at the feedback on each side, and compares what it costs to practise writing seriously with each.

The plain test: what a student actually gets

Eighty-six percent against seventy-eight. On 36 essays, that is three more correct scores for Prepamigo than for GPT 5.6 sol, six more than GPT 5.6 luna, and ten more than GPT-4o mini. Put another way, Prepamigo makes 37% fewer scoring mistakes than GPT 5.6 sol, 55% fewer than luna and 67% fewer than GPT-4o mini.

What the GPT models were given matters. Each was told it was an experienced CELPIP writing examiner, given the task exactly as the candidate saw it, including the required points for the email task and the options for the survey task, and given the essay. It was asked for one score from 0 to 12. That is what you would type into ChatGPT, and it is a fairer test than a bare “grade this” because the model at least knows what the task asked for.

Prepamigo’s scorer received the same task and the same essay, plus what it always has in production: real graded sample essays for that task at each level to compare against, and a rule layer that checks the things a rater checks first, such as whether every required point is covered, whether a survey response actually picks a side, whether an email has a greeting and a closing, and whether the response is anywhere near the length the task asks for.

Where the eight points come from: level by level

Weak essays (verified score 4 or 5)

12 official essays. The one level where the GPT models edge ahead.

83%

GPT 5.6 sol

83%

GPT 5.6 luna

83%

GPT-4o mini

75%

Prepamigo Writing Scorer

On weak essays, all three GPT models are at 83% and Prepamigo is at 75%. This is the one level where ChatGPT is ahead, and we report it as such. Prepamigo’s three misses all went one notch too high, to a 6 or 7, on essays that were short but tidy. GPT 5.6 sol’s two misses were both a 6. Weak essays are the easy end of the scale for a chatbot: short, generic and error-prone writing looks weak to anyone.

Middle essays (verified score 7 or 8)

12 official essays. The level most candidates are at.

100%

Prepamigo Writing Scorer

75%

GPT 5.6 sol

67%

GPT-4o mini

58%

GPT 5.6 luna

In the middle, Prepamigo scored all twelve essays inside the 7 to 8 band. GPT 5.6 sol scored nine; its three misses were all a 6, one notch under. GPT 5.6 luna scored seven, pulling three essays down to a 6, one to a 5 and one up to a 9. GPT-4o mini scored eight and pushed the other four up to a 9.

A 7 or 8 is the most common real score on the test, and it is the score that decides whether a candidate has reached their target. GPT 5.6 sol will tell one such candidate in four that they have slipped to a 6. Luna will tell four in ten. Neither is a disaster in the way Claude’s middle-band results are, but it is the difference between a practice score you can plan around and one you have to second-guess.

Strong essays (verified score 10 or above)

12 official essays. Misses here are a 9: the essay was strong and the grader held back.

83%

Prepamigo Writing Scorer

75%

GPT 5.6 sol

67%

GPT 5.6 luna

25%

GPT-4o mini

At the top, Prepamigo recognised ten of twelve strong essays, GPT 5.6 sol nine, luna eight, and GPT-4o mini three. Every miss but one was a 9. GPT-4o mini gave a 9 to nine of the twelve strong essays: it simply does not hand out a 10. If your writing is strong and you check it with the free tier of ChatGPT, which still routes a lot of traffic to small models, you will be told you are a 9 almost every time.

A 9 from ChatGPT on an essay you thought was excellent is more likely to be a held-back 10 than a true 9. Rewrite it once with the model answer beside it and score it again; a consistent 10 across two attempts is a 10.

Why a chatbot slips in the middle and at the top

A middle-level CELPIP essay covers the task, is organised, and is also generic: reasons asserted rather than developed, safe vocabulary, sentences of one shape. A strong one is not flawless; it has a slip or two, and a plain sentence among the developed ones. A model with no graded examples to compare against has to decide from impression. GPT 5.6 sol’s impression is a little severe: the generic 7 reads as a 6, the strong-with-a-slip 10 reads as a 9. Luna is more severe still. GPT-4o mini treats 9 as the ceiling.

Prepamigo’s scorer does not judge an essay against an idea of good writing. It compares it with real graded essays for the same task, one at each level, all of them real responses with real slips. A generic but complete essay lands next to the middle sample because that is what it resembles. A strong essay with one clumsy sentence lands next to the strong sample, which also has one. On top of that comparison sits a rule layer for the things a model is bad at counting: whether every required point is there, whether a survey response picked a side, whether the email has a greeting and a closing, and how long it is. A missed required point holds task fulfillment to 8 or below regardless of the English, because that is how the test works.

The workflow gap: what you have to bring

For writing, ChatGPT is easy to use as a grader, and we will not pretend otherwise. Paste the essay, get a number, ask for an explanation. No recording, no transcript. For a quick second opinion on a single essay it is convenient and, with GPT 5.6 sol, reasonably accurate.

What you do not get is the task. ChatGPT can invent a CELPIP-style prompt if you ask, but it is guessing at the required points, the survey options and the word count the real test uses. You do not get a required-point check, because it does not know which points the real test would list. You do not get a model answer rewritten from your own essay at the right length. And each essay you paste counts against your message allowance.

On Prepamigo, you open a task from the bank, the timer runs, you write, and about a minute after you submit you have four dimension scores, an overall score, a list of anything you missed, corrections quoted from your own text, and a rewritten model answer. The fortieth essay of the month costs you the same effort as the first.

The same essay, the same score

After the headline run we scored all 36 essays twice more on different days. Prepamigo’s writing scorer scored 86%, 89% and 86% across the three runs, with the same twelve middle-level essays inside the band every time. GPT 5.6 sol scored 78%, 83% and 81%: better on its second and third runs, and a reminder that a single chatbot answer carries a few points of luck in either direction. If you do grade with ChatGPT, ask twice and keep the lower answer.

Feedback: explanation versus evidence

ChatGPT explains well. Ask why an essay is a 6 and it will tell you, in clear English, with suggestions. The catch is that the explanation is written to justify a number already chosen, so when the number is a notch low, the explanation finds faults to match. A candidate who was really at a 7 reads a page about weaknesses that were not decisive.

Prepamigo’s feedback works from the text up. Every correction quotes the exact words from your essay, and anything the checker cannot find in your text is dropped before you see it. A missed required point is named. Each dimension gets one limiting factor, or a plain statement that nothing is holding it back. And you get a model answer rewritten from your own response at the length the task asks for. The section below lists what changed.

What is new in your writing feedback

The score is only half of what you get back. The writing feedback layer was rebuilt alongside the scorer, and everything in it is anchored to your own text. Here is what changed.

  • Corrections you can find in your own essay. Every grammar, spelling and vocabulary correction quotes the exact words from your response, so the app can highlight them where you wrote them. Anything the checker cannot find word for word in your essay is dropped before you see it.
  • A model answer, built from your answer. You get a rewritten version of your own response that keeps your ideas, covers every required point and lands in the 150 to 200 words the task asks for. If the first rewrite misses that length, it is redone once before it is shown.
  • The required point you missed, by name. For the email task, the feedback lists the bullet points from the prompt that your response did not cover. A missed point also holds the task-fulfillment score to 8 or below, so the number and the feedback never disagree.
  • One limiting factor per dimension. For each of the four dimensions, the feedback names the single thing holding that score down, in concrete terms, or says plainly that nothing is and what is carrying it. Deciding a dimension has no problem is a real answer, not a gap.
  • Criticism in the right place. A weak tone is a task-fulfillment problem, a vague word choice is a vocabulary problem, a run-on sentence is a readability problem. Each point is filed under the dimension that owns it instead of being lumped into the overall summary.
  • Sentence-level rewrites. Specific sentences from your essay come back with an improved version and one line on why it is better, so you can see the change rather than read about it.
  • Your strongest and weakest dimension, side by side. The feedback tells you which of the four dimensions is carrying your score and which is holding it back, so you know what to practise next without decoding four numbers.

Every quotation in the feedback is checked against your essay before it is displayed. If the checker cannot find the words, the line is removed. That is why the feedback never tells you to fix something you did not write.

What it costs to practise every day

ChatGPT Plus is US$20 a month, about CA$28, billed monthly, as published in September 2026, with a cap on messages per rolling window; ChatGPT Pro, at US$200 a month, about CA$276 before tax, raises the cap. The free tier routes much of its traffic to smaller models, and on this test the smallest of them recognised three strong essays in twelve. None of the plans include a task bank, a timer, a required-point check, a model answer, or a score calibrated against real graded essays.

Prepamigo is CA$24.98 a month, or CA$8.25 a month on the annual plan, which is CA$98.98 for the year, tax included, and you can cancel at any time. Every plan includes unlimited scoring with no daily, session or weekly cap, unlimited attempts, and the full question bank: 80 complete practice sets and more than 2,000 practice tasks across Writing, Speaking, Reading and Listening, written to follow the task types, timing and format of the CELPIP General test.

When ChatGPT is the better tool

  • A quick second opinion on one essay. GPT 5.6 sol is a decent grader and takes ten seconds. If you already have the task and just want another view, it is fine.
  • Rewriting a sentence five ways. Asking for alternatives to a clumsy sentence is quick and the suggestions are good.
  • Talking it through. If you want to argue about why a reason is weak, a conversation is the right shape. Prepamigo gives you a verdict and evidence; it does not chat.
  • Anything outside the test. ChatGPT is a general assistant and Prepamigo is not.

What it should not be is the only thing that tells a 7 or 8 writer where they stand, or the thing that tells a strong writer they are a 9. Those are the two places it is wrong most often.

FAQs

Can ChatGPT score my CELPIP writing? Yes, and better than most chatbots: GPT 5.6 sol was right 78% of the time on 36 official responses, luna 69%, GPT-4o mini 58%, against 86% for Prepamigo.

Is Prepamigo more accurate than ChatGPT for writing? By eight points against the best GPT model. Middle-level essays 100% against 75%; strong essays 83% against 75%; weak essays 75% against 83%.

Which GPT model should I use? GPT 5.6 sol. Not GPT-4o mini, which gave a 9 to nine of twelve strong essays.

How much does each cost? ChatGPT Pro is US$200 a month, about CA$276 before tax; Plus is US$20 with message caps. Prepamigo is CA$24.98 a month tax included, or CA$8.25 a month on the annual plan, with unlimited scoring and 80 practice sets.

For the same comparison against Claude, read Prepamigo vs Claude for writing. For how the writing scorer works, read inside the Prepamigo writing scorer. Or write a response and see the score.

Prepamigo vs ChatGPT for CELPIP Writing: Accuracy, Limits and Cost | PrepAmigo