Lot Machine: Multimodal Lot Extraction from Auction Catalogs
2026-08-31 • Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionArtificial IntelligenceComputation and LanguageDigital Libraries
AI summaryⓘ
The authors worked on a way to automatically pull detailed information about auctioned items from old German auction catalogs. They tested different smart computer models that read images and text together, checking how well these models follow rules to produce clean, structured data. They looked at options from big companies and smaller, local setups to see what works best under real-world limits like budget and privacy. Their results show that with some human help, this method can help researchers study lots of historical auction info more easily.
Provenance ResearchAuction CatalogsVision-Language ModelsStructured Data ExtractionJSON FormatCultural HeritageMachine-readable DataPrompt StrategiesData PrivacyQuantized Models
Authors
Mathias Zinnen, Alisha Mund, Sabine Lang, Lukas Hüttner, Thomas Gorges, Vincent Christlein
Abstract
For provenance research and art market studies, auction catalogs are an essential resource to trace specific objects over time and space. While historical auction catalogs follow established domain conventions, their internal formatting remains highly variable, and their large-scale analysis is currently restricted by the lack of machine-readable representations of the auction lots. We propose a pipeline to automatically extract structured lot-level metadata from German Sales, a large database of historical auction and sales catalogs from the 19th and 20th centuries. Using a manually annotated test set of representative catalog pages, we evaluate Vision-Language Models (VLMs) under varying prompt strategies and constrained decoding frameworks. To reflect the practical constraints faced by cultural heritage institutions, including budget, compute resources, and data privacy requirements, we benchmark the methods across different deployment modes ranging from commercial providers to locally hosted, quantized models. We find that commercial endpoints establish the performance ceiling, while institutional gateways offer a viable, privacy-preserving alternative. Local deployments remain feasible, but strictly require enforcing the output structure during generation to guarantee a valid JSON format. While varying degrees of human-in-the-loop correction are still necessary, this work demonstrates that a VLM-based pipeline can successfully unlock historical auction catalogs for large-scale automated analysis.