Summary
When police officers respond to a cybercrime quickly, small mistakes can cause big problems in catching criminals. The authors looked at different AI tools that might help these officers make better decisions during that critical first hour. They found that some AI methods work better than others, but current AI tools still have risks and don’t always follow legal rules perfectly. The study also shows that existing ways to test these AI tools don’t check if they handle evidence safely or answer simple questions well enough. The authors suggest making new tests focused on protecting evidence and making sure AI advice fits legal needs.
What this means in practice
- •For cybersecurity response teams: Use findings to prioritize retrieval-augmented AI systems that better fit the needs and skills of frontline responders for early cyber incident actions.
- •For legal technology developers: Design evaluation benchmarks that emphasize robustness to naive queries and strict evidence handling to improve AI tools used in judicial cybercrime processes.
A survey. It maps existing work.
Abstract
The actions of frontline law enforcement officers in the initial hour of a cyber incident play a vital role in determining the ultimate success of an investigation. The minor mistakes they commit might result in irreversible critical impacts. The integrity of the investigation can be compromised, and the prosecution of cyber criminals can be hindered due to minor mistakes that happen in the initial hour. These are mainly because of the volatile nature of digital artifacts that might lead to procedural errors and evidence attrition. This paper provides a systematic survey of decision-support architectures designed to assist first responders of a cybercrime, categorizing them into playbooks, Large Language Models (LLMs), Retrieval-Augmented Generation (RAG) frameworks, and Agentic AI systems. The survey critically considers the constraints of limited technical proficiency and inconsistent forensic infrastructure in a practical scenario. Our analysis identifies RAG-based systems as a relatively viable intermediate solution due to their natural language adaptability. However, significant risk factors like prompt sensitivity and the potential for confident hallucinations in legal contexts pose a major challenge. Furthermore, we review current benchmarks in cybersecurity and demonstrate that they are not sufficient to capture the specific safety and legal requirements of law enforcement, focusing on the initial hour of the cybercrime. We conclude by arguing for the necessity of a new evaluation benchmark focused on naive query robustness and evidence preservation, so as to ensure that AI-driven guidance aligns with the mandatory demands of judicial proceedings.