Applications of Risk Science to AI Fairness Evaluation: Principles, Challenges, and Best Practices

Computers and SocietyArtificial Intelligence

Summary

The authors looked at how research on AI fairness and bias in hiring compares to established ideas in risk science, which studies how to understand and manage risks. They found that most studies talk about how serious unfairness might be but don't clearly show how uncertain those estimates are. They then showed how to include risk science ideas in testing an AI system that screens resumes. Finally, they proposed a tool called the AI Risk Report Card to help explain AI risks clearly to decision-makers. Their work suggests combining risk science and AI evaluation could improve how we study and talk about the social effects of AI.

Authors

Kyra Wilson, Sabrina Kang, Saloni Dash, Aylin Caliskan

Abstract

Scholarly work which aims to describe potential societal impacts (e.g., risks) of proliferating technology (especially related to artificial intelligence or other algorithmic systems) is likely to have an impact beyond the scientific communities it was written for, given that general society itself is a primary object of study. However, it is an open question whether the current practices of AI evaluation scholarship follow the principles and best practices established by risk science, which aims to systematically generate knowledge related to understanding, assessing, communicating, managing, and governing risk. In this work, we examine this in depth by conducting a literature review of scholarly works purporting to evaluate the bias or fairness of technological systems used for tasks related to hiring and employment. Through analysis of 22 common fairness evaluation metrics and studies using them, we find that most characterize the severity of bias- or fairness-related consequences but do not follow best practices to characterize the uncertainty around either the occurrence of these consequences or severity estimates. Next, we conduct a case study of fairness evaluation for an AI-mediated resume screening task and demonstrate how principles of risk science can be incorporated into such an evaluation. Finally, we propose the AI Risk Report Card, which facilitates the reporting and communication of risk assessment results to stakeholders in positions to act based on the predicted risks. The outcomes of these activities suggest that further research at the convergence of risk science and AI evaluation can lead to advancements in AI assessments of societal impact by enabling shared frameworks to evaluate and discuss AI risks both within and outside of the scientific community.