Papers for

website designers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Multimodal simulation improves website user experience recommendations

Automatic multimodal UX improvement recommendations from LLM agent user simulations

Abstract: Evaluating user experience (UX) on live websites through user testing is expensive, subjective, and difficult to scale. LLM agents offer a promising route to automating UX testing by simulating realistic user behaviour. However, existing simulation approaches typically lack multimodality and require time-consuming manual review to extract actionable insights. We formalise UX improvement recommendation from simulation data as a structured natural language generation and ranking problem, and establish an evaluation protocol using expert annotation and LLM-as-a-Judge. We present AMUSER, a multimodal framework which simulates user behaviour and automatically generates prioritised UX improvement recommendations from resulting data. We evaluate AMUSER on commercial websites and show that its recommendations substantially outperform those from text-only simulation (NDCG@3 = 0.758 versus 0.359) at an 89% lower simulation cost. Our results suggest an asymmetric role of multimodality: visual access during simulation improves recommendations through richer traces, while providing visual inputs during recommendation generation can modestly degrade quality. We also discuss practical deployment lessons from applying AMUSER to commercial websites.

Sat 19 SeptComputation and LanguageArtificial IntelligenceHuman-Computer Interaction
The gist
Testing how easy and pleasant websites are to use is usually expensive and slow because it needs real people. The authors created a system called AMUSER that uses AI to simulate users interacting with websites, including both what they see and do. This lets the system suggest ways to improve websites automatically and faster, with better suggestions than previous approaches that looked only at text actions. Interestingly, giving the AI visual information during the testing part helps a lot, but giving visuals when making suggestions can slightly hurt quality.
Open → 2609.22971v1