Evaluating Large Language Models on Greek Primary School Mathematics Problems: Accuracy, Reasoning, and Educational Implications
Φόρτωση...
Αρχεία
Ημερομηνία
2025-12-12
Συγγραφείς
Volakakis, Argyris
Tzivinikou , Sotiria
Τίτλος Εφημερίδας
Περιοδικό ISSN
Τίτλος τόμου
Εκδότης
Παραπομπή
Παραπομπή
Άδεια Creative Commons
Εκτός εάν σημειώνεται διαφορετικά, η άδεια αυτού του αντικειμένου περιγράφεται ως Attribution-NonCommercial-NoDerivatives 4.0 International
Περίληψη
Abstract — The rapid advancement of Large Language Models (LLMs) such as the GPT family is reshaping expectations for automated tutoring, formative feedback, and assessment in education. Despite impressive progress in natural language understanding, the ability of these models to perform reliable mathematical reasoning, especially in non-English languages and primary-level educational contexts, remains insufficiently examined. This study investigates how accurately and consistently current LLMs solve mathematical problems from the Greek primary school curriculum and how their reasoning patterns relate to task difficulty.
A dataset of 70 mathematics problems was created based on the 3rd- and 5th-grade Greek mathematics textbooks. The problems are categorized by grade, topic (e.g, number sequences, fractions, geometry), and a difficulty scale (1–4), which is validated by three experienced educators and two LLMs. Using a Python-based evaluation pipeline, each problem was submitted in the Greek language to four widely known OpenAI models, i.e., GPT-3.5-turbo, GPT-4o, and GPT-4.1-mini, and GPT-5-mini, under two prompting conditions: (i) the LLM to provide a direct answer and (ii) the LLM to provide the complete step-by-step reasoning.
All answers were forwarded to a dedicated Python module to assess accuracy based on the corresponding ground truth in the database, as well as time and cost spent per model. Additionally, step-by-step responses were analyzed using an LLM-as-a-Judge approach, where the GPT-5 model rated the reasoning quality, completeness, intermediate calculations, understanding, and consistency of each model, besides accuracy. This dual analysis offers an initial exploration of hybrid methods for assessing AI reasoning in educational tasks.
The evaluation revealed systematic performance differences across grade levels, mathematical domains, and difficulty levels. The overall accuracy is higher on 3rd-grade problems and computationally oriented topics, reaching approximately 90%, while it decreased substantially for problems with greater difficulty or more complex reasoning. As
task difficulty increased, accuracy declined by about 8% for GPT-5-mini and by more than 40% for GPT-3.5-turbo, indicating a sensitivity to cognitive demand, particularly for earlier-generation models.
The contribution of this work lies in showing where and why these differences in LLM performance occur and how they relate to specific task types. Beyond the quantitative findings, the study considers key pedagogical and ethical issues associated with deploying LLMs in primary mathematics education. This research proposes a small-scale yet reproducible framework for evaluating LLMs’ mathematical reasoning in non-English elementary settings. Future work will expand the dataset, examine additional LLMs beyond OpenAI’s models, and evaluate the quality of their tutoring capabilities in mathematics. These analyses provide initial evidence relevant to the design of LLM-based tutoring tools for elementary mathematics. The findings aim to inform both developers designing educational AI systems and educators seeking to integrate such technologies responsibly into classroom practice.
Περίληψη
Περιγραφή
Λέξεις-κλειδιά
Large Language Models, Mathematics Education, Elementary School, AI in Education, Automated Assessment

