YOMI-Bench: A Benchmark for Evaluating Kanji Reading and Phonological Understanding of LLMs for Japanese

2026-07-01Computation and Language

Computation and Language
AI summary

The authors created YOMI-Bench, a test to see how well large language models (LLMs) can read and understand kanji, the complex characters used in Japanese. Since one kanji can have many possible readings, it's tricky for models to pick the right one just from the text. They tested various LLMs, including some made just for Japanese and some commercial ones, and found that all of them struggled with reading kanji correctly. This shows that current models still have a hard time with this important part of Japanese language.

kanjilarge language modelsJapanese languagephonological understandingmultilingual modelsJapanese-specific modelsbenchmarklanguage generationreading tasksmodel evaluation
Authors
Ryota Mibayashi, Hiroya Takamura, Hitomi Yanaka
Abstract
We propose YOMI-Bench, a benchmark for evaluating kanji reading and phonological understanding of large language models (LLMs) for Japanese. In Japanese, a single kanji character often has multiple possible readings, making it difficult to infer the correct reading from surface-level text alone. Due to these linguistic characteristics, it is empirically known that LLMs exhibit low performance in kanji reading for Japanese. The proposed YOMI-Bench consists of four tasks specifically designed to evaluate kanji reading performance in Japanese. In our evaluation using YOMI-Bench, we assessed one multilingual open LLM, four Japanese-specific open LLMs, and five commercial LLMs. As a result, we found that even Japanese-specific models show low performance, and that commercial models also perform poorly on generation tasks that require consideration of kanji readings.