WeaveData improves multimodal analysis with self-correcting AI plans

WeaveData: A Multimodal Data Analysis System with Self-Critiquing and Self-Evolving LLM Plans

DatabasesArtificial Intelligence

Summary

Multimodal data analysis means answering questions about tables, text, and images all together, which is tricky. The authors created WeaveData, a system that helps AI language models plan these analyses better by checking its own steps and fixing mistakes. WeaveData also learns from past errors and talks with users to clear up confusing questions. This makes analyzing complex combined data more reliable and interactive.

What this means in practice

  • For data analysis teams: Use WeaveData to build automated systems that answer complex questions over combined tables, text, and images with fewer errors and clearer explanations.
  • For business intelligence developers: Integrate WeaveData to enhance multimodal query handling by diagnosing plan failures and refining analysis workflows interactively with users.

Authors

Min Jia, Shihao Zhou, Jun-Peng Zhu, Peng Cai, Kai Xu, Chao Zhang, Li Li, Aoying Zhou, Heng Long, Qiu Cui, Liu Tang, Qi Liu

Abstract

Multimodal data analysis, which answers questions over relational tables, text, and images, has attracted growing attention in the data management community. Large language models (LLMs) enable such analysis in natural language by generating analysis plans over relational and semantic operators. However, LLM-generated plans are error-prone: a plan may silently compute something other than what was asked, fail during execution, or return a result that misses the question. This paper presents WeaveData, a multimodal data analysis system with self-critiquing and self-evolving LLM plans. First, WeaveData generates a typed logical plan for each question and critiques it step by step before execution, and it checks the executed result against the question afterwards. Second, WeaveData evolves a plan that fails or misses the question: it diagnoses the failure with the actual data, reuses the results that remain valid, and accumulates planning experience for later questions. Third, WeaveData grounds planning in a metadata knowledge graph of all modalities, clarifies ambiguous questions with the user, and backs every model judgment with evidence in an interactive notebook. We demonstrate WeaveData on two public multimodal datasets.