Models show gaps in judging and fixing user interface aesthetics

AUV-Bench: Aesthetic Understanding and Generation Evaluation for User Interfaces

Artificial Intelligence

Summary

User interface design involves judging how nice a page looks and fixing problems when needed. The authors studied how well current AI models can score, diagnose, and fix UI designs compared to professional designers. They found that while models can somewhat agree on overall looks, they often struggle to correctly identify specific design issues and rarely fix them well. This shows that current AI models only partially understand UI aesthetics and need better connections between spotting problems and making improvements.

What this means in practice

  • For ui designers: Assess AI tools for aesthetic scoring and diagnosis to support professional UI design workflows.
  • For software developers: Incorporate findings to improve AI-assisted interfaces that generate or repair web page designs automatically.

Authors

Zhijie Deng, Ling Li, Junhao Ji, Siwei Lyu, Zhipeng Xu, Zulong Chen, Rongyao Fang, Shuai Bai, Xuming Hu, Jiaheng Wei

Abstract

Multimodal foundation models are increasingly used for evaluating and generating user interfaces (UIs), often producing seemingly reasonable aesthetic judgments and visually plausible pages. However, under professional design scrutiny, their behavior can differ substantially from that of human designers. In professional design practice, designers rely on a systematic set of aesthetic principles that consistently guide judgment, diagnosis, repair, and creation. A coherent aesthetic capability should therefore connect aesthetic judgment with design actions. Existing evaluations, however, typically assess these abilities in isolation, making it difficult to determine whether task-level success reflects a shared aesthetic understanding or merely fragmented task-specific competence. To address this gap, we introduce AUV-Bench, developed in collaboration with professional UI designers around 1,395 executable web interfaces and four tasks: aesthetic scoring, diagnosis, repair, and text-to-UI generation. The tasks share a pool of UIs and aesthetic principles, with diagnosis and repair further aligned on 660 controlled-degradation instances to enable instance-level analysis of judgment and action. Evaluation of 12 models reveals a capability imbalance: models show moderate agreement with professional designers in holistic aesthetic scoring, yet exact diagnosis-chain success peaks at only 24.7%. On the aligned diagnosis-repair cases, correct judgments and successful repairs do not consistently coincide, exposing a Judgment-Action Gap between identifying aesthetic problems and successfully acting on them. In open-ended generation, even leading models achieve only moderate aesthetic quality under human-calibrated evaluation. Overall, current models exhibit partial aesthetic competence, but still lack the fine-grained understanding and judgment-action coherence required for reliable UI design.