Measures for competent generative AI use at work show weak correlation
Beyond AI Literacy: A Structured Review and Exploratory Meta-Analysis of Measures for Competent Generative-AI Use
Human-Computer InteractionArtificial Intelligence
Summary
It is hard to measure how well people use AI tools at work because different kinds of tests, like self-ratings and performance tasks, don’t strongly agree with each other. The authors reviewed studies and found no clear way to replace actual performance tests with self-reports. They looked at tools that test knowledge and how people check or trust AI, but none cover all necessary skills. They propose a new testing approach but haven’t yet shown it works better.
What this means in practice
- •For human resources teams: Develop better workplace tests to assess employees’ competent use of generative AI tools based on knowledge and trust factors.
- •For compliance officers: Use layered assessment methods to evaluate controls and oversight of AI tool use, helping monitor safe and effective practice.
A survey. It maps existing work.
Authors
Daniele Veri'
Abstract
Researchers assessing competent generative-AI use at work must choose among self-reports, objective tests, and measures of oversight and reliance. We conducted a structured, seeded review of 24 focal empirical publications, starting from the 2024 COSMIN-based review and adding a targeted update through 17 August 2026. We grouped the measures into four domains: knowledge and use, epistemic oversight, reliance calibration, and operational control of tool-using agents. In an exploratory meta-analysis, we pooled three direct subjective-objective correlations from one research program (REML r = .055; Hartung-Knapp 95% CI [-.047, .156]; combined reported N = 2,765). We could not resolve a discrepancy between the largest study's reported correlation and p-value, leaving its weight uncertain. Adding a synthetic mean of 12 cross-factor correlations from a fourth study gave r = .079 (95% CI [-.025, .181]). This sensitivity analysis concerns a broader comparison. From this small evidence base, we cannot establish a population correlation, validate workplace cutoffs, or justify substituting self-ratings for performance scores. We identified tests of foundation knowledge (AICOS-S and GLAT) and measures of verification, reliance, trust, and dependency. We found no validated individual-level instrument in the focal corpus that tests the full combination of agent scope, permissions, recovery, state isolation, independent review, and evidence-based closure; some cover subsets. We propose a four-layer workplace battery with non-compensatory decision rules, but have not tested its thresholds or whether it improves on other assessment approaches.