Last updated: August 16, 2026
Short answer: no test-prep AI system, including ChatGPT, is accurate enough on its own to trust with a Writing band score that decides your visa or university offer â and the research backs that up. At English Learning Point (ELP), every single Writing task is marked by a real, qualified human examiner, band by band, criterion by criterion. No AI shortcuts. Here’s the evidence for why that matters, and exactly how ELP’s marking works.
What the research actually says about AI scoring accuracy
The most directly relevant study is Osama Koraishi’s 2024 paper in Language Teaching Research Quarterly, “The Intersection of AI and Language Assessment: A Study on the Reliability of ChatGPT in Grading IELTS Writing Task 2.” Koraishi took 55 real IELTS Writing Task 2 essays already graded by official IELTS assessors, then had ChatGPT-4 grade the same essays independently, and compared the two sets of scores statistically.
- Intraclass Correlation Coefficient (ICC): 0.814 (95% CI: 0.702â0.887) â a measure of how consistently the two raters (ChatGPT and the human examiners) agreed on the same essays.
- Cohen’s Weighted Kappa: 0.811 (95% CI: 0.726â0.896) â a stricter agreement measure that corrects for chance agreement.
- Mean official grade: 6.03 (SD 1.14) vs. mean ChatGPT grade with SD 1.09 â close on average, Wilcoxon test p = 0.91 (no statistically significant difference in the average).
On paper, an ICC and Kappa both above 0.8 sounds reassuring â statisticians generally call that “substantial agreement.” But Koraishi’s own conclusion is the part that matters for you as a test-taker: ChatGPT showed real, individual-essay discrepancies significant enough that the study explicitly recommends human oversight before any high-stakes decision is made from an AI score. “Substantial agreement” on average across 55 essays is not the same as “reliable” for the one essay that decides your band score. A confidence interval that runs from 0.70 to 0.89 also means that on some essays, agreement was meaningfully weaker than the headline number suggests â and you have no way of knowing in advance whether your essay will be one of them.
In plain terms: AI scoring tools are a genuinely useful practice aid â fast, free or cheap, and good for catching obvious grammar and structure issues. They are not a substitute for a trained human examiner when the score has real consequences.
AI scoring vs. human marking, side by side
| Factor | AI Scoring Tools | ELP Human Marking |
|---|---|---|
| Agreement with official examiners | ICC 0.814 / Kappa 0.811 â “substantial” on average, with real per-essay outliers (Koraishi, 2024) | Marked directly against the official IELTS/PTE/TOEFL/OET band descriptors by a human examiner |
| Understands context, nuance, argument quality | Limited â pattern-matches vocabulary, structure, and surface coherence | Full â evaluates whether your argument actually answers the question and holds together logically |
| Catches Task Achievement / Task Response gaps | Inconsistent â this is the criterion AI struggles with most | Central focus â this is often where students lose the most marks unknowingly |
| Personalised, actionable feedback | Generic suggestions, same phrasing across many users | Written specifically for your essay, your recurring errors, your target band |
| Explains why a mark was given | Rarely, or in vague generalities | Always â band-by-band breakdown against the 4 official criteria |
| Accountability if you disagree | None â no one to ask | Direct access to the examiner (Mudasser) to ask questions |
How ELP’s human marking actually works
Every Writing task submitted at English Learning Point is read and marked personally â never auto-graded, never passed through an AI scoring API. Marking follows the same four official band criteria examiners use in the real exam (Task Achievement/Response, Coherence and Cohesion, Lexical Resource, Grammatical Range and Accuracy), with a band awarded for each, plus a written explanation of exactly what pulled the score up or down. Where a script is borderline between two bands, it’s marked conservatively against the stricter interpretation â the same discipline the real exam applies â so students get an honest, exam-realistic number rather than an inflated one.
Why this matters more in Pakistan specifically
For most ELP students, a Writing band score isn’t an academic exercise â it’s the number standing between them and a visa, a nursing registration abroad, or a university seat. Pakistan sends some of the highest volumes of IELTS and OET candidates in the region, and the gap between a 6.5 and a 7.0 on Writing can mean the difference between qualifying for a preferred visa category or not. In a market where free AI scoring tools are heavily marketed as “instant band predictors,” it’s worth being clear-eyed about what the actual research shows: those tools are a reasonable way to practise, but they are not accurate enough, essay by essay, to plan a visa application or a university deadline around. A wrong AI-predicted band that turns out to be a full band different from the real exam can cost a student a missed intake, a re-sit fee, or a delayed move â all avoidable with an honest, human-checked estimate first.
Who’s marking your Writing
Mudasser Iqbal, founder of English Learning Point, personally marks every student’s Writing task, band by band. He holds an MA in English and an MBA (HRM), is currently pursuing CELTA certification, and has around 15 years of teaching experience preparing Pakistani students for IELTS, PTE, TOEFL, and OET. That combination â a real examiner’s eye plus a training background, applied consistently to every single submission â is the thing no AI scoring tool currently replicates.
Frequently asked questions
Is AI scoring accurate enough to trust for a real IELTS band prediction?
Not reliably enough for a high-stakes decision. The best available research (Koraishi, 2024) found “substantial” but imperfect agreement between ChatGPT and official IELTS examiners (ICC 0.814, Kappa 0.811) across 55 essays â with real per-essay discrepancies the study itself flags as needing human oversight. Use AI tools for quick practice feedback, not for planning a visa or university deadline.
Does ELP use any AI to mark Writing tasks?
No. Every Writing task submitted through ELP is read and marked personally, band by band, against the official criteria â never auto-scored.
Who marks my Writing task at ELP?
Mudasser Iqbal, ELP’s founder â MA English, MBA (HRM), pursuing CELTA, ~15 years of teaching experience â personally marks every submission.
How long does human marking take compared to an AI tool?
An AI tool returns a score in seconds; ELP’s human marking takes longer because it’s a real, considered read of your essay â but you get an accurate, explained band score rather than an instant guess. Turnaround details are confirmed when you register for a batch.
Can I see exactly why I got the band score I did?
Yes â every marked script comes with a band-by-band breakdown against the four official criteria (Task Achievement/Response, Coherence and Cohesion, Lexical Resource, Grammatical Range and Accuracy), so you know exactly what to fix.
Is human marking available for PTE, TOEFL, and OET too, or just IELTS?
All four. Every writing/essay component across IELTS, PTE, TOEFL, and OET at ELP is human-marked, not AI-scored.
What if I disagree with my score?
You can ask directly â because a real person marked your script, there’s someone to explain the reasoning, unlike an AI tool with no one behind the score.
Get your Writing marked by a real examiner, not an algorithm
If you want an honest, accurate band score you can actually plan your visa or university timeline around, message ELP on WhatsApp and get started.