Large vision language models tested on art for education
MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education
Artificial IntelligenceComputation and LanguageComputer Vision and Pattern Recognition
Summary
Many advanced AI systems can understand pictures and text together, but we don't know how well they work with art in classrooms. The authors created MUSE, a test that checks how well AI can understand artistic images, including cultural and emotional meanings, especially those related to Southeast Asia and Western art. They tested different AI models and found that they struggle most with feelings and complex reasoning about images. This benchmark helps improve AI tools that support learning with art and diverse cultures.
What this means in practice
- •For educational technology developers: Evaluate and improve AI tools that interpret artistic images for language learning and cultural education.
- •For digital museum software teams: Assess AI models for enriching visitor interaction with art by understanding its cultural and emotional context.
Authors
Luyao Zhu, Xun Wei Yee, Wei Li, Mun Thye Mak, Wee Siong Ng
Abstract
Large vision-language models have achieved remarkable progress in multi-modal understanding, yet their capabilities in educational settings remain insufficiently evaluated. In AI-assisted language learning, models must interpret artistic imagery, understand its semantic, affective, and cultural content, and reason about visual context to support meaningful interaction. However, existing benchmarks primarily focus on real-world images or domain-specific educational reasoning, providing limited coverage of artistic educational content. To address this gap, we introduce MUSE, a benchmark for evaluating large vision-language models on artistic image understanding in situated educational applications. MUSE decouples image annotation from question generation, enabling diverse tasks with controllable difficulty while reducing annotation effort. It comprises twelve tasks spanning visual perception, semantic and affective interpretation, culture understanding, and compositional reasoning, together with diverse artistic images deliberately curated to center Singaporean and Southeast Asian multicultural contexts alongside Western art traditions, covering multiple themes and difficulty levels. Evaluation of open-source and proprietary models reveals substantial disparities across capability dimensions, particularly in affective interpretation and compositional reasoning. Our analysis further identifies common failure modes and key challenges for developing trustworthy multi-modal models for education. We hope MUSE will serve as a standardized benchmark for advancing multi-modal understanding in situated educational applications.