2026 IJRSE – Volume 15 Issue 19
Available Online: 16 August 2026
Author/s:
Koskinas, Emmanouil
Democritus University of Thrace, Greece (koskinas.manolis@gmail.com)
Dosi, Ifigeneia
Democritus University of Thrace, Greece (idosi@hs.duth.gr)
Chadjipapa, Elina*
Democritus University of Thrace, Greece (elina@hotmail.com)
Gavriilidou, Zoe
Democritus University of Thrace, Greece (zoegab@otenet.gr)
Abstract:
This study examined the agreement between ChatGPT-4 Plus and experienced human raters in the assessment of argumentative essays written by Greek upper secondary school students. Specifically, it investigated (a) the degree of agreement between ChatGPT and human raters and (b) the relative contribution of content, text organization, and language (vocabulary and grammar) to overall writing scores. The dataset comprised 267 argumentative essays that had previously been evaluated by trained educators using the official Greek Ministry of Education scoring rubric. The same essays were subsequently assessed by ChatGPT-4 Plus using standardized prompting procedures. Results indicated moderate to good agreement between ChatGPT and human raters for both overall and trait-level scores, with agreement comparable to that observed between the two human raters. No significant differences were found in overall scores. At the trait level, however, ChatGPT assigned significantly higher scores for content, whereas human raters assigned higher scores for language; no significant differences emerged for text organization. Regression analyses showed that all three scoring components contributed significantly and relatively evenly to overall scores in both scoring systems, although ChatGPT demonstrated stronger intercorrelations among the scoring dimensions. Overall, the findings suggest that ChatGPT-4 Plus shows considerable promise as an automated essay scoring tool in Greek secondary education. However, systematic differences in evaluative emphasis indicate that it is currently best viewed as a complementary tool rather than a replacement for expert human judgment, particularly in high-stakes writing assessment.
Keywords: automated essay scoring, assessment tool, human raters, ChatGPT, Greek Upper Secondary Education
DOI: https://doi.org/10.5861/ijrse.2026.26368
Cite this article:
Koskinas, E., Dosi, I., Chadjipapa, E., & Gavriilidou, Z. (2026). Evaluating ChatGPT-4 as an automated essay scoring tool: A comparative study with human raters in Greek upper secondary education. International Journal of Research Studies in Education, 15(19), 183-198. https://doi.org/10.5861/ijrse.2026.26368
* Corresponding Author
