MineCEraft: Evaluating Language Models as Construction Engineers in the World of Minecraft

Artificial Intelligence

Summary

The authors created MineCEraft, a new tool to test how well large language models (LLMs) can follow construction instructions in Minecraft. It includes 723 expert-written tasks with clear ways to check if the instructions are done right, covering 17 different building challenges. Using this tool, the authors tested current LLMs and found common mistakes and difficulties these AI models face when doing construction tasks. Their work helps understand where LLMs succeed or struggle in simulated building projects.

Authors

Sewoong Lee, Risham Sidhu, Julia Hockenmaier, Yoonhwa Jung

Abstract

We introduce MineCEraft (Minecraft Construction Engineering Benchmark, pronounced mine-see-ee-raft), an easy-to-use, open-source benchmark designed to systematically evaluate the reliability and limitations of LLMs for construction tasks in Minecraft. The MineCEraft benchmark comprises 723 domain-expert hand-crafted natural-language instructions with programmatically verifiable evaluation, spanning 17 distinct task categories, providing a safe and controllable experimental environment for assessing LLMs' ability to perform realistic construction engineering tasks. With this benchmark, we conduct an in-depth evaluation of state-of-the-art LLMs and perform a detailed error analysis, revealing key failure modes and practical challenges in applying LLMs to construction engineering tasks.