How Useful are LLMs for Grammar Engineering? Cantonese ParGram Resources and Controlled Experimental Evaluation with English Baselines
2026-08-24 • Computation and Language
Computation and Language
AI summaryⓘ
The authors created new linguistic resources for Cantonese and tested large language models (LLMs) like GPT-5.4 and gpt-oss-120b to see if they could generate formal grammar rules from sentences. They found that GPT-5.4 performed better, especially when given direct formal structures rather than just sentences. However, both models had trouble managing complex grammar rules that interact in multiple ways. The study shows that while LLMs can help with some parts of making grammars, human experts are still needed to check and improve the results.
Cantonese ParGramlarge language models (LLMs)grammar engineeringphrase-structure ruleslexical entriesformal grammarmachine-processable grammarsprompting conditionsAI-assisted workflowslinguistic expertise
Authors
Chit-Fung Lam
Abstract
This paper presents new Cantonese ParGram resources and evaluates LLMs for knowledge-driven grammar engineering within a controlled experimental paradigm. Using Cantonese ParGram resources as gold standards, with corresponding English baselines, we investigate whether OpenAI's gpt-oss-120b and GPT-5.4 can generate machine-processable grammars from sentences and target formal structures under systematically varied prompting conditions. GPT-5.4 outperformed gpt-oss-120b, while grammars generated from target formal structures generally outperformed those generated from sentences. Although both models could generate locally plausible phrase-structure rules, lexical entries, and templates, they often struggled to coordinate interacting formal constraints, especially in multi-construction settings. The results characterize both the capabilities and limitations of current LLMs for potential integration into AI-assisted expert workflows: LLMs may support intermediate stages of grammar development, but human linguistic expertise remains central to analysis, validation, and refinement. The study also contributes new Cantonese symbolic grammatical resources.