MELLON - Multimodal Enhanced LLM for Online Navigation

2026-08-10Artificial Intelligence

Artificial Intelligence
AI summary

The authors studied how computer programs that browse websites can do better by understanding both text and images together. They focused on a test called WebShop and created three new tools to help these programs think and plan using multimodal inputs. One tool, called MELLON, made a noticeable improvement in how well the programs complete tasks after a short training. The authors suggest that using and improving multimodal methods is important for making web navigation smarter.

web navigation agentsmultimodal reasoningWebShop benchmarkLLM (Large Language Model)task completion accuracyimage-text alignmentmultimodal inputsonline navigationtraining epochsplanning abilities
Authors
Ruiyu Li, Haoyang Cai, Zhitong Guo, Tong Hu
Abstract
Web navigation agents are capable of addressing various types of tasks on different websites. Current baselines on web navigation are either unimodal or lack strong reasoning abilities given multimodal inputs. Focusing on the WebShop benchmark, a real-world website simulation, we explore the alignment of text and images, as well as multimodal reasoning and planning abilities, to enhance the performance of web navigation agents. We propose three innovative multimodal enhancements: Multimodal Enhanced LLM for Online Navigation (MELLON), VQAgent, and Multimodal Ranker. MELLON demonstrates a significant improvement in task completion accuracy, with a 9.26% increase after just one epoch of training. Our findings suggest the necessity of further exploration into multimodal approaches, with a focus on more extensive training and alignment strategies to enhance the effectiveness of web navigation agents.