Study · September 2026
How accurate is our AI IELTS Writing scoring?
We ran our AI scorer over 55 Writing Task 2 essays, twice each, and compared its bands with reference bands. Here is what we found, including where it still gets things wrong and why these numbers are not an examiner comparison.
By Oleksii Vasylenko · Evaluation run 27 September 2026
Limitations
Read this before the numbers
No examiner-marked ground truth yet
None of the reference bands in this study were given by a certified IELTS examiner. Until they are, these figures describe agreement with our references, not agreement with an official test result.
Stored bands are not examiner marks
45 of the essays are real submissions. Their reference is the band our own system stored when the essay was submitted, so for these essays the study measures consistency with our earlier scoring, not correctness.
Small sample
55 essays, one task type (Writing Task 2). The 10 authored essays only cover bands 6.0–8.0, and some band groups are too small to report on their own.
Practice estimates, not official scores
Our bands are practice estimates to guide study. They are not IELTS results and are not endorsed by the British Council, IDP or Cambridge.
What we measured
The evaluation harness sends each essay to the same scoring prompts production uses and records the overall band (the rounded mean of the four criteria: Task Response, Coherence and Cohesion, Lexical Resource, Grammatical Range and Accuracy). Each essay was scored twice so we could check consistency.
We tested both scorers: the full evaluation signed-in users get, and the quick score from the free essay checker.
45
Real submissions
Reference: the band our system stored at submission. Bands 1.0–8.5, a few per band.
10
Authored essays
Written in-house to a target band (6.0–8.0). The target is our judgement, not an examiner's.
×2
Runs per essay
Temperature 0, reasoning off.
How to read the metrics
- Accuracy figures use the first run of each essay that returned a score.
- Low-band inflation counts weak essays (reference 4.5 or below) scored a full band or more too high — the error that most misleads a learner.
- Privacy. We publish aggregates only: no essay text, no identifiers, and no group with fewer than 20 essays.
Results: before and after the prompt change
“Before” is the scoring prompt production used until 27 September 2026; “after” is the rewritten prompt now in production. Both were run with the same model and settings, so the difference comes from the prompt.
| Metric | Before | After |
|---|---|---|
| Average error (MAE)Mean absolute distance from the reference band. Lower is better. | 0.84 | 0.52 |
| Within ±0.5 bandShare of essays scored no more than half a band from the reference. | 55% | 70% |
| Average biasPositive = scored higher than the reference, negative = lower. | −0.39 | −0.41 |
| Average error, authored essaysThe 10 essays written to a known band. | 0.75 | 0.35 |
| Weak essays inflated by ≥1 bandReference band 4.5 or lower, scored a full band or more too high. | 6 of 20 | 1 of 20 |
| Same band on both runsShare of essays that got an identical overall band when scored twice. | 90% | 67% |
| Failed responsesCalls that returned no usable score (broken or incomplete output). | 4 of 110 | 1 of 110 |
| Metric | Before | After |
|---|---|---|
| Average error (MAE)Mean absolute distance from the reference band. Lower is better. | 0.62 | 0.51 |
| Within ±0.5 bandShare of essays scored no more than half a band from the reference. | 64% | 73% |
| Average biasPositive = scored higher than the reference, negative = lower. | −0.27 | −0.16 |
| Average error, authored essaysThe 10 essays written to a known band. | 0.40 | 0.30 |
| Weak essays inflated by ≥1 bandReference band 4.5 or lower, scored a full band or more too high. | 4 of 21 | 2 of 21 |
| Same band on both runsShare of essays that got an identical overall band when scored twice. | 93% | 67% |
| Failed responsesCalls that returned no usable score (broken or incomplete output). | 0 of 110 | 0 of 110 |
Across the 45 real submissions alone, the full evaluation’s average error went from 0.86 to 0.56. The authored-essay figures rest on only 10 essays, so treat them as a direction, not a precise number.
Where it is still wrong: bias by band
The main remaining error is at the top: essays with a reference of band 7 or above were scored −0.62 band on average by the full evaluation. Strong writers should expect our estimate to be conservative.
Weak essays are no longer inflated much: 1 of 20 came out a full band or more too high, and the average bias for that group is −0.35. The middle group (n = 14) is below our 20-essay reporting threshold, so it is counted in the overall figures but not shown separately.
| Reference band | Bias before | Bias after |
|---|---|---|
| 1–4.5 | +0.33 (n = 20) | −0.35 (n = 20) |
| 5–6.5 | not reported (n = 13) | not reported (n = 14) |
| 7 and above | −1.23 (n = 20) | −0.62 (n = 20) |
| Reference band | Bias before | Bias after |
|---|---|---|
| 1–4.5 | +0.02 (n = 21) | +0.10 (n = 21) |
| 5–6.5 | not reported (n = 14) | not reported (n = 14) |
| 7 and above | −0.87 (n = 20) | −0.67 (n = 20) |
Consistency and failures
The new prompt is more accurate but less repeatable. With the old prompt, 90% of essays got the same band on both runs of the full evaluation; with the new one, 67%. When the two runs differed it was usually by half a band near a boundary; the largest gap we saw was 1.0 band.
Failed responses — where no usable score came back — fell from 4 to 1 of 110 calls on the full evaluation. The quick score had 0 failures in 110 calls. The app checks every response against the same rules and does not show a band from one that fails them.
What changed on 27 September 2026
The rewritten Task 2 prompt:
- Uses one shared rubric distilled from the public Task 2 band descriptors, with the details that separate bands 4 to 8 on each criterion.
- Asks for evidence from the essay before each criterion score.
- Drops numeric bands from its worked example, so the model is not anchored to them.
- Removes an old instruction that set a minimum score.
- Adds hard limits for under-length, off-topic or memorised, and non-English answers.
- Runs at temperature 0, so the same essay tends to get the same score.
Writing Task 1 scoring was not changed and is not covered by this study.
Other settings we tested
Before choosing the settings above, we ran the old prompt with other models and settings. Turning on the model’s reasoning mode broke many responses at our output budget. The larger model had the lowest average error on the full evaluation, but it failed more often and inflated weak essays far more on the quick score, so we did not adopt it. We have not yet tested it with the new prompt.
| Scorer · model · reasoning · temp. | Failed | Avg error | ±0.5 | Weak inflated |
|---|---|---|---|---|
| Full · deepseek-flash · low · 0.3 | 96/110 | 0.94 | 38% | 0/6 |
| Full · deepseek-flash · off · 0.3 | 2/110 | 0.86 | 48% | 8/20 |
| Full · deepseek-flash · off · 0.7 | 1/110 | 0.70 | 61% | 5/20 |
| Full · deepseek-v4-pro · off · 0.3 | 7/110 | 0.45 | 77% | 5/20 |
| Quick · deepseek-flash · low · 0.3 | 51/110 | 0.62 | 62% | 6/14 |
| Quick · deepseek-flash · off · 0 | 0/110 | 0.61 | 62% | 6/21 |
| Quick · deepseek-flash · off · 0.3 | 0/110 | 0.62 | 62% | 4/21 |
| Quick · deepseek-v4-pro · low · 0.3 | 25/110 | 0.73 | 61% | 14/19 |
| Quick · deepseek-v4-pro · off · 0.3 | 0/110 | 1.01 | 38% | 18/21 |
Accuracy columns only count calls that returned a score, so settings with many failures are measured on fewer essays.
How we'll improve this
The biggest gap in this study is the reference itself. The next version needs essays marked blind by certified IELTS examiners and experienced IELTS teachers, so we can measure our scores against human marks instead of our own earlier output.
If you are a certified examiner or an IELTS teacher and would mark a small set of anonymised Task 2 essays without seeing our scores, please email us. We will publish the results here, including where the AI disagrees with you, and credit contributors who want to be named.
No form, no sign-up — just an email.
Cite this study
IELTS.international (2026). How accurate is our AI IELTS Writing scoring? An evaluation of AI band scores on 55 IELTS Writing Task 2 essays, September 2026. https://www.ielts.international/research/ai-scoring-accuracy