AI tool combines secret industry archives for reliable analysis
INDRA: A New AI Tool for Exploring Tobacco, Fossil Fuel, and Chemical Industry Archives
Digital LibrariesComputation and LanguageComputers and Society
Summary
Many secret documents from tobacco, fossil fuel, and chemical companies have been hard to study using large language models (LLMs) because they were never organized for such tools. The authors created INDRA, a platform that collects these archives into one place and uses special rules to keep the AI strictly inside the selected documents. INDRA also shows where information comes from and separates facts from guesses, helping users check the AI's answers instead of just trusting them. This makes it easier to do wide-ranging research on hidden industry practices while avoiding mistakes common in other AI tools.
What this means in practice
- •For policy analysts: Search combined industry archives to build evidence-based regulatory reports without mixing in outside information.
- •For legal teams: Use a unified AI platform to prepare litigation strategies by reliably referencing internal industry documents with traceable sources.
Authors
Daniel Akselrad, Robert N. Proctor
Abstract
Five decades of litigation have disgorged hundreds of millions of pages of formerly secret business records from the tobacco industry, along with documents from the makers of drugs, chemicals, food, firearms, and fossil fuels. Yet these archives have been effectively inaccessible to general-purpose large language models (LLMs) because they have never been compiled into an LLM-readable corpus. Chatbots may be familiar with some of the materials contained in such archives but, with no direct access to the documents, they are vulnerable to hallucination and other defects. Here we introduce INDRA, a research platform designed to remedy such failures by embedding the conventions of archival historiography into a system-level protocol governing every output. The platform federates UCSF's Industry Documents Library, Columbia and CUNY's ToxicDocs, Stanford's SRITA, and other heretofore siloed collections, and provides three interlinked safeguards: (1) a closed evidentiary sandbox confines the model to a user-selected corpus, blocking retrieval from external sources that could introduce bias; (2) real-time provenance tagging marks the boundary between archival evidence and parametric inference; and (3) a system-level protocol enforced by deterministic scripts guides the structure of every output. Together these safeguards prevent the model from conflating "the documents say X" with "I think X" or "I learned X from prior training." The result is an LLM-powered research partner enabling massive multi-archival investigations, a tool whose outputs are designed to be checked rather than trusted, and whose architecture makes the conditions of knowledge production visible and auditable. Three case studies demonstrate the method's analytical value and limitations, including what we call the Heraclitus effect, the steppingstone dilemma, and the gullibility (or mafia) problem.