Indonesian culture digital library uses AI to improve data collection
Artificial Intelligence-Assisted Digital Inventory of Cultural Heritage & Traditional Knowledge: Case for Indonesian Open Digital Library of Culture
Artificial IntelligenceDigital LibrariesHuman-Computer InteractionInformation Retrieval
Summary
Collecting cultural knowledge from many sources is difficult because information is scattered, mixed with errors, or incomplete. The authors describe a step-by-step AI process to find, check, and add cultural data from the web to Indonesia’s digital library. Their approach carefully combines automated methods with important checks by humans to ensure accuracy and respect cultural meaning. They also make sure that human contributions are never overwritten by machines. The paper discusses how this balances technology with ethical and cultural sensitivity.
digital librarycultural heritageartificial intelligenceweb crawlingdata extractionmultilingual processingknowledge representationbayesian fusionhuman-machine interactiondata quality
Authors
Hokky Situngkir
Abstract
The Indonesian Digital Library of Culture (Perpustakaan Digital Budaya Indonesia, PDBI; budaya-indonesia.org) is a participatory platform that has collected tens of thousands of entries on Nusantara cultural heritage through public contribution since 2007. Manual contribution faces three structural barriers: coverage (knowledge is scattered across languages and sites), integrity (open sources mix authentic documentation with noise), and completeness (subjects are recorded but their data remain shallow). This paper presents a methodological framework for autonomous, AI-based harvesting of cultural knowledge from the open web, designed to expand corpus coverage while intensifying per-entry data depth. The methodology is organised as a five-stage economic funnel: focused crawling, multilingual extraction and canonicalisation, vector encoding with blocking, agentic decision-making, and idempotent publication, under the principle of deterministic orchestration, agentic decisions. Each stage is formalised: funnel economics and optimal filter ordering; crawl-frontier dynamics as a subcritical branching process that explains the necessity of recurrent re-seeding; fact-level novelty via a containment measure; Bayesian multi-source evidence fusion with elevated publication thresholds for sacred categories; exactly-once effects via idempotent upserts and the transactional outbox; sliding-window inference budgeting with a reservation protocol; statistical quality auditing; and seed selection as submodular coverage maximisation. The framework retains four high-value human roles: curator of direction, escalation approver, quality auditor, and guardian of meaning, while machine autonomy is raised in stages. Ethical, legal, and cultural-sensitivity implications are discussed, including the architectural guarantee that the machine never overwrites human contributions.