WuYu-EnvLE-Bench: A Benchmark for Evaluating Large Language Models in Environmental Law Enforcement

2026-07-20Artificial Intelligence

Artificial Intelligence
AI summary

The authors created WuYu-EnvLE-Bench, a test set using real environmental enforcement cases to see how well large language models (LLMs) can help with enforcement decisions. They found that while LLMs do well on tasks with clear rules, they struggle with complex tasks like linking evidence, spotting contradictions, combining information from different sources, and making procedural judgments. Bigger models don’t always perform much better, especially on tasks that need careful reasoning. The authors suggest that better methods are needed for LLMs to handle evidence and rules effectively in environmental enforcement.

Large Language ModelsEnvironmental EnforcementBenchmarkEvidence-chain ConstructionContradiction DetectionRegulatory StandardsModel ScalingRule-bounded TasksProcedural JudgmentMulti-source Integration
Authors
Ziliang Yang, Yi Zhang, Kaijun Lin, Jiachao Ke, Haihong Xu, Zongguo Wen
Abstract
Large language models (LLMs) are increasingly considered for environmental enforcement, but their ability to produce traceable enforcement decisions remains unclear. We introduce WuYu-EnvLE-Bench, a benchmark built from real enforcement cases, regulatory standards, and expert review. It contains 2,521 benchmark instances, 14 tasks, and 12 pollution-medium subdomains across pre-enforcement, in-enforcement, and post-enforcement workflows. Using Absolute Environmental Enforcement Score (AES) and Intelligent Enforcement Index (IEI), we evaluate open-source and closed-source LLMs across capability, response quality, and resource efficiency. Results show that LLMs perform well on rule-bounded tasks but remain unreliable in evidence-chain construction, contradiction detection, multi-source integration, and procedural judgment. Model scaling also shows diminishing returns: medium-sized models approach leading models in structured tasks, while larger models do not reliably overcome evidence-reasoning bottlenecks. WuYu-EnvLE-Bench highlights the need for evidence-grounded, rule-aware, and task-adaptive enforcement reasoning.