IELTS.international

Study · September 2026

How accurate is our AI IELTS Writing scoring?

We ran our AI scorer over 55 Writing Task 2 essays, twice each, and compared its bands with reference bands. Here is what we found, including where it still gets things wrong and why these numbers are not an examiner comparison.

By Oleksii Vasylenko · Evaluation run 27 September 2026

Limitations

Read this before the numbers

  • No examiner-marked ground truth yet

    None of the reference bands in this study were given by a certified IELTS examiner. Until they are, these figures describe agreement with our references, not agreement with an official test result.

  • Stored bands are not examiner marks

    45 of the essays are real submissions. Their reference is the band our own system stored when the essay was submitted, so for these essays the study measures consistency with our earlier scoring, not correctness.

  • Small sample

    55 essays, one task type (Writing Task 2). The 10 authored essays only cover bands 6.0–8.0, and some band groups are too small to report on their own.

  • Practice estimates, not official scores

    Our bands are practice estimates to guide study. They are not IELTS results and are not endorsed by the British Council, IDP or Cambridge.

What we measured

The evaluation harness sends each essay to the same scoring prompts production uses and records the overall band (the rounded mean of the four criteria: Task Response, Coherence and Cohesion, Lexical Resource, Grammatical Range and Accuracy). Each essay was scored twice so we could check consistency.

We tested both scorers: the full evaluation signed-in users get, and the quick score from the free essay checker.

45

Real submissions

Reference: the band our system stored at submission. Bands 1.0–8.5, a few per band.

10

Authored essays

Written in-house to a target band (6.0–8.0). The target is our judgement, not an examiner's.

×2

Runs per essay

Temperature 0, reasoning off.

How to read the metrics

  • Accuracy figures use the first run of each essay that returned a score.
  • Low-band inflation counts weak essays (reference 4.5 or below) scored a full band or more too high — the error that most misleads a learner.
  • Privacy. We publish aggregates only: no essay text, no identifiers, and no group with fewer than 20 essays.

Results: before and after the prompt change

“Before” is the scoring prompt production used until 27 September 2026; “after” is the rewritten prompt now in production. Both were run with the same model and settings, so the difference comes from the prompt.

Full evaluation
MetricBeforeAfter
Average error (MAE)Mean absolute distance from the reference band. Lower is better.0.840.52
Within ±0.5 bandShare of essays scored no more than half a band from the reference.55%70%
Average biasPositive = scored higher than the reference, negative = lower.−0.39−0.41
Average error, authored essaysThe 10 essays written to a known band.0.750.35
Weak essays inflated by ≥1 bandReference band 4.5 or lower, scored a full band or more too high.6 of 201 of 20
Same band on both runsShare of essays that got an identical overall band when scored twice.90%67%
Failed responsesCalls that returned no usable score (broken or incomplete output).4 of 1101 of 110
Quick score (free essay checker)
MetricBeforeAfter
Average error (MAE)Mean absolute distance from the reference band. Lower is better.0.620.51
Within ±0.5 bandShare of essays scored no more than half a band from the reference.64%73%
Average biasPositive = scored higher than the reference, negative = lower.−0.27−0.16
Average error, authored essaysThe 10 essays written to a known band.0.400.30
Weak essays inflated by ≥1 bandReference band 4.5 or lower, scored a full band or more too high.4 of 212 of 21
Same band on both runsShare of essays that got an identical overall band when scored twice.93%67%
Failed responsesCalls that returned no usable score (broken or incomplete output).0 of 1100 of 110

Across the 45 real submissions alone, the full evaluation’s average error went from 0.86 to 0.56. The authored-essay figures rest on only 10 essays, so treat them as a direction, not a precise number.

Where it is still wrong: bias by band

The main remaining error is at the top: essays with a reference of band 7 or above were scored −0.62 band on average by the full evaluation. Strong writers should expect our estimate to be conservative.

Weak essays are no longer inflated much: 1 of 20 came out a full band or more too high, and the average bias for that group is −0.35. The middle group (n = 14) is below our 20-essay reporting threshold, so it is counted in the overall figures but not shown separately.

Full evaluation
Reference bandBias beforeBias after
1–4.5+0.33 (n = 20)−0.35 (n = 20)
5–6.5not reported (n = 13)not reported (n = 14)
7 and above−1.23 (n = 20)−0.62 (n = 20)
Quick score (free essay checker)
Reference bandBias beforeBias after
1–4.5+0.02 (n = 21)+0.10 (n = 21)
5–6.5not reported (n = 14)not reported (n = 14)
7 and above−0.87 (n = 20)−0.67 (n = 20)

Consistency and failures

The new prompt is more accurate but less repeatable. With the old prompt, 90% of essays got the same band on both runs of the full evaluation; with the new one, 67%. When the two runs differed it was usually by half a band near a boundary; the largest gap we saw was 1.0 band.

Failed responses — where no usable score came back — fell from 4 to 1 of 110 calls on the full evaluation. The quick score had 0 failures in 110 calls. The app checks every response against the same rules and does not show a band from one that fails them.

What changed on 27 September 2026

The rewritten Task 2 prompt:

  • Uses one shared rubric distilled from the public Task 2 band descriptors, with the details that separate bands 4 to 8 on each criterion.
  • Asks for evidence from the essay before each criterion score.
  • Drops numeric bands from its worked example, so the model is not anchored to them.
  • Removes an old instruction that set a minimum score.
  • Adds hard limits for under-length, off-topic or memorised, and non-English answers.
  • Runs at temperature 0, so the same essay tends to get the same score.

Writing Task 1 scoring was not changed and is not covered by this study.

Other settings we tested

Before choosing the settings above, we ran the old prompt with other models and settings. Turning on the model’s reasoning mode broke many responses at our output budget. The larger model had the lowest average error on the full evaluation, but it failed more often and inflated weak essays far more on the quick score, so we did not adopt it. We have not yet tested it with the new prompt.

Other model settings tested with the old prompt
Scorer · model · reasoning · temp.FailedAvg error±0.5Weak inflated
Full · deepseek-flash · low · 0.396/1100.9438%0/6
Full · deepseek-flash · off · 0.32/1100.8648%8/20
Full · deepseek-flash · off · 0.71/1100.7061%5/20
Full · deepseek-v4-pro · off · 0.37/1100.4577%5/20
Quick · deepseek-flash · low · 0.351/1100.6262%6/14
Quick · deepseek-flash · off · 00/1100.6162%6/21
Quick · deepseek-flash · off · 0.30/1100.6262%4/21
Quick · deepseek-v4-pro · low · 0.325/1100.7361%14/19
Quick · deepseek-v4-pro · off · 0.30/1101.0138%18/21

Accuracy columns only count calls that returned a score, so settings with many failures are measured on fewer essays.

How we'll improve this

The biggest gap in this study is the reference itself. The next version needs essays marked blind by certified IELTS examiners and experienced IELTS teachers, so we can measure our scores against human marks instead of our own earlier output.

If you are a certified examiner or an IELTS teacher and would mark a small set of anonymised Task 2 essays without seeing our scores, please email us. We will publish the results here, including where the AI disagrees with you, and credit contributors who want to be named.

Email hello@ielts.international

No form, no sign-up — just an email.

Cite this study

IELTS.international (2026). How accurate is our AI IELTS Writing scoring? An evaluation of AI band scores on 55 IELTS Writing Task 2 essays, September 2026. https://www.ielts.international/research/ai-scoring-accuracy
Please link to this page when you quote a figure, and mention that the references are not examiner marks. Numbers may be updated when we re-run the study; the evaluation date is part of the citation.