COMPARATIVE EVALUATION OF THREE LLMS FOR CEFR-ALIGNED READING PASSAGE AND ITEM GENERATION
DOI:
https://doi.org/10.20319/ictel.2026.221222Keywords:
Large Language Models, Automatic Item Generation, Reading Comprehension, CEFR, Higher-Order SkillsAbstract
Research Objectives: This study compares three large language models (Deepseek, GPT-4, Baidu Ernie 4.0) in automatically generating CEFR A1–C2 reading comprehension passages and test items.
Methodology: Each model produced 72 passages (narrative, expository, argumentative, instructional) and 120 four-option multiple-choice items covering nine CEFR skill descriptors. Five certified language assessment experts independently rated all materials on nine quality dimensions using 5-point Likert scales. Inter-rater reliability was excellent (Fleiss’ κ = 0.82–0.91). Data were analysed with Kruskal-Wallis H tests and post-hoc Dunn–Bonferroni corrections.
Findings: Passage and item quality were strong up to B2 level (M > 4.2/5.0 across models), but declined notably at C1–C2 (M = 3.08–3.78). Higher-order skills (e.g., rhetorical purpose evaluation, implicit attitude inference) scored significantly lower (p < .001). Deepseek outperformed GPT-4 on CEFR alignment (p = .002) yet remained 0.5–0.7 points below human-authored benchmarks.
Research Outcomes: Current LLMs can reliably generate psychometrically sound reading materials up to B2 level, but remain inadequate for fully automated high-stakes C1–C2 assessment without intensive human post-editing.
Future Scope: Enhance LLM capabilities for rhetorical complexity, cohesion manipulation, and higher-order item design targeting advanced proficiency levels.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.
Copyright of Published Articles
Author(s) retain the article copyright and publishing rights without any restrictions.
This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.

