AI haiku mimic human poets but are hard to tell apart
Authorship attribution and aesthetic evaluation of AI poetry: a case study with Haiku
Computation and LanguageArtificial Intelligence
Summary
This paper looks at how well computers can write traditional Japanese haiku poems and if people can tell these computer poems from human ones. The authors tested multiple AI language models and asked Japanese students to guess which poems were written by AI or humans. Results showed that some AI poems were so good that people could not reliably distinguish them from human poems. Interestingly, poems that seemed more fluent and poetic were often thought to be human-made, even when they were not. This suggests that as AI improves, spotting AI-generated poetry may become increasingly difficult.
What this means in practice
- •For creative writing software developers: Improve AI-based poetry generators to create more human-like haiku for creative writing tools.
- •For content moderation teams: Develop better methods to identify AI-generated poetry within constrained forms where detection is currently unreliable.
Authors
Livia Oddi, Simone Scardapane, Toru Sugimoto, Donatella Genovese
Abstract
This paper investigates the generation and human evaluation of Japanese haiku by contemporary Large Language Models (LLMs), focusing on authorship perception and aesthetic judgment within a constrained poetic form. Using a few-shot prompting strategy, Japanese haiku were generated across a heterogeneous set of large language models, including open- and closed-source systems, medium-scale and large-scale architectures, models with native or adapted Japanese support, and multilingual proprietary models. These AI-generated haiku were combined with human-written ones and presented in a questionnaire distributed to students at Japanese universities in Tokyo. The survey assessed whether respondents could distinguish between AI-generated and human-written haiku and which cues informed their judgments. Recognition accuracy varied across models. GPT-5, Gemini 2.5, and StableLM-7B performed at approximately chance level (approx 0.50), whereas LLM-JP, Gemma-2B, and LLaMA-2 showed moderate detectability (approx 0.59-0.67). However, recognition was strongly item-dependent. Ratings of fluency, coherence, poeticness, and related aesthetic dimensions predicted perceived humanness but not correct classification, indicating an attribution bias linked to aesthetic evaluation and revealing a dissociation between aesthetic evaluation and true authorship detection. The extended analysis additionally examines generation-constraint adherence, participant-level characteristics, and exploratory LLM-based evaluations of haiku authorship. Overall, the findings suggest that as LLMs improve, surface-level creative plausibility may reduce reliable human discrimination within constrained poetic settings.