Is the ACL Responsible NLP Checklist a Box-Ticking Exercise? A Large-Scale Analysis of EMNLP 2025

2026-08-10Computation and Language

Computation and Language
AI summary

The authors studied the use of EMNLP 2025 Responsible NLP Checklists, which are meant to encourage transparency, ethics, and awareness of societal impacts in NLP research. They collected and analyzed over 73,000 checklist responses from papers and found common problems, such as authors treating ethics as an afterthought and providing weak or contradictory answers. Many authors appeared to downplay risks or social impacts of their work. The authors also identified design flaws in the checklist and suggested improvements like requiring longer explanations and more careful consideration of potential risks in future versions.

Responsible NLPEthics checklistTransparencyEMNLPResearch ethicsSocietal impactChecklist designNLP research practicesAuthor complianceEvaluation metrics
Authors
Nusrath Jinnath, Wei Zhao
Abstract
Responsible NLP practice includes a) transparency, b) ethics, and c) societal impacts. The Responsible NLP Checklist aims to push these goals, and promote responsible practice. Recently, ACL released the EMNLP 2025 Checklists to aid transparency on the current research practice, which we focus on. We curate and release the first two datasets of: a) all the checklist responses and justifications from the EMNLP 2025 Main and Finding tracks; b) checklist reference linking to paper sections. We also provide the first analysis of recent EMNLP Checklists, by examining $73,922$ responses and justifications to them. For the Main track, we find that authors isolate ethics questions of the Checklist from the paper's bulk, mimicking the trend of ethics being an afterthought. We then examine \texttt{NO} responses. We find $44.9\%$ of justifications are poor or bad-faith, being brief or empty. Then, we find significant issues with the checklist design and effort of authors, namely that $6\%$ of all checklists contained logical contradictions between parent and child responses. We also find evidence of surface compliance for responsible ethics, with $53\%$ authors dismissing potential risks or social impacts of their work, for which there should be none. We compare this to the Findings track, noticing a similar trend in both tracks. Lastly, we discuss the implications of the checklist design and provide recommendations for future checklist iterations. Including: a) enforcing a minimum word count, b) enforcing more scrutiny on the risks of appliances.