Greek language models tested on national exam questions reveal strengths and limits
Pass or Fail? Evaluating LLMs on Two Greek Examination Benchmarks
Computation and Language
Summary
Large Language Models (LLMs) that understand Greek are tested with real exam questions used for Greek schools and universities. The authors created two new test sets from these exams to see how well different Greek and multilingual LLMs perform. They found that a smaller, specially adapted Greek model does very well in language tasks compared to bigger models. Also, they learned that traditional ways to judge answers don’t work well for complex reasoning, and that giving extra example questions sometimes helps but can overwhelm smaller models.
What this means in practice
- •For natural language processing engineers: Design specialized Greek language models that perform well on humanities tasks despite smaller size by adapting models linguistically.
- •For educational assessment developers: Use realistic Greek exam benchmarks to evaluate AI systems’ abilities across diverse academic subjects and question types.
Authors
Panagiota Kyriazi, Eleni Kasoura, Prokopis Prokopidis
Abstract
The rapid advancement of Large Language Models (LLMs) imposes a thorough evaluation of their linguistic and analytical capabilities as well as constraints, particularly for a language with limited benchmark coverage such as Greek. To address the limited availability of comprehensive benchmarks in this domain, we introduce Prot-Ex and Pan-Ex, two benchmarks consisting of questions from entrance exams for Greek Model and Experimental schools as well as the Panhellenic exams (the Greek national university entrance examinations). These benchmarks are employed to assess the performance of text-only LLMs-including the Greek-adapted KriKri-8B-Instruct, Llama-3.1-8B, Gemma-4-26B, and Qwen-3-32B-across diverse academic disciplines (Modern Greek, Mathematics, Physics, etc.) and task formats (closed, structured, and open-ended), including textualized visual context (i.e., image descriptions). Our findings indicate the localized KriKri-8B significantly outperforms its base model, successfully rivalling much larger LLMs in linguistically demanding humanities tasks. By leveraging an LLM-as-a-Judge methodology, we expose the inadequacy of traditional lexical metrics for evaluating complex reasoning. Crucially, we uncover a few-shot prompting paradox: while synthetic examples improve accuracy in closed-ended questions, they severely overload the context window of 8B models in structured tasks, causing significant performance degradation. Ultimately, this study suggests targeted linguistic adaptation offsets lower parameter counts in specialized domains, despite the fragility of smaller models to prompt verbosity.