Frontier vision-language models have overtaken young adults at detecting AI-generated portraits -- but not their calibration

2026-08-31Human-Computer Interaction

Human-Computer InteractionComputer Vision and Pattern Recognition
AI summary

The authors tested 19 vision-language AI models to see how well they can spot AI-generated face portraits compared to real photos. They found that in mid-2026, several new models outperformed young adults in identifying fake images, with the best model detecting AI images correctly over 90% of the time. However, these models tend to be biased in their judgment, unlike humans who maintain a more balanced approach between trust and suspicion. The study also showed that changing example images used for training still caused models to make errors about 25% of the time. Overall, while AI models are getting better at recognizing fakes, they do not yet match humans in subtle judgment balance.

vision-language modelsAI image generationface portraitsmodel accuracyhuman calibrationbalanced accuracyimage detectionbias in AIsensitivityd-prime
Authors
Sunwhi Kim, Sunyul Kim, Meounggun Jo, Jini Tae
Abstract
AI image generators now create face portraits that are hard to tell from real photographs. Vision-language models (VLMs) are increasingly proposed to flag such images. We benchmarked 19 VLMs on the same 198 face portraits -- real photographs and identity-matched ChatGPT-4o and Imagen 3 versions -- under the same task as our earlier study of 1,667 adults (85% correct overall; accuracy fell steeply with age). The June-2026 cohort of 14 models only matched adults in their 20s-30s. Four weeks later the ceiling broke. Among five July-2026 releases under the identical protocol, gpt-5.6-sol reached 92.8% balanced accuracy (five-draw mean 92.1%), clearly above adults in their 20s (88.5%), and claude-fable-5 detected every AI image while averaging 91.9%. Model sensitivity now exceeds young adults decisively (d' up to 3.4 versus ~ 2.4). What has not been overtaken is human calibration. Model criteria spread from c = -1.10 to +1.45 while humans sit near zero at every age; both new leaders are biased (+0.44, -0.97), and only a few mid-ranked models approach the human balance. Changing the labelled examples still flipped about one answer in four. The best machines now out-see young adults here, without matching the human balance between suspicion and trust.