Coherence aware metric improves evaluation of generated text quality

Coherence-Aware Distributional Evaluation of Open-Ended Text Generation

Computation and LanguageArtificial Intelligence

Summary

Current ways to measure how well computers write text often miss if a story or explanation makes sense overall. The authors created a new method called CHORD that checks if generated text is consistent and connected throughout, not just word by word. They do this by looking inside a language model's understanding of the text and comparing it to human writing. CHORD is better at spotting when the text goes off-topic or contradicts itself, helping to judge computer-written text more like a human would.

What this means in practice

  • For natural language processing engineers: Evaluate and rank text generation models by coherence sensitivity beyond token-level measures during model development and fine-tuning.
  • For content moderation teams: Identify generated text outputs that contain contradictions or incoherent passages that simpler metrics fail to flag, improving quality control.

Authors

Jinnuo Liu, Junhao Zhu, Weifeng Jiang, Haoming Liu, Hongyi Wen

Abstract

Existing metrics for open-ended text generation measure likelihood, lexical diversity, or distributional similarity in generic representation space, yet they can miss fundamental dimensions of quality. A prominent blind spot is global coherence: a generated passage may be locally fluent while remaining globally contradictory, causally inconsistent, or topically disconnected. Such failures can still preserve the token-level and lexical statistics that existing metrics rely on. We identify representation as a central bottleneck in detecting these failures and introduce CHORD (Coherence-aware Hidden-state Open-generation Reference Distance), a coherence-sensitive distributional metric. CHORD encodes generated and human-written corpora in the hidden-state space of a frozen LLM using a coherence-eliciting prompt, and compares the resulting distributions using MMD with an RBF kernel. To validate that the metric responds to coherence degradation but not generic textual change, we construct a counterfactual evaluation suite that pairs graded coherence-degrading perturbations with meaning-preserving controls. CHORD selectively detects relation, discourse, structural, and mixture failures that perplexity, entropy, MAUVE, FBD, and MMD-based baselines either miss or cannot separate from benign rewriting. Factorial ablations show that representation is the primary source of coherence sensitivity,while RBF-MMD improves sample efficiency once the relevant distinctions become visible. Larger backbones capture finer-grained distinctions, but coherence prompting improves selectivity only when the backbone can follow the prompt.On unconditional generation and prefix continuation, CHORD yields model rankings that strongly align with human judgments of whether outputs make sense and appear human-written. Together, these results establish representation design as central to reliable distributional evaluation.