Visible but Not Yet Curatable: Characterizing the Curatability of Compact and Derived Open LLM Artifacts

Human-Computer InteractionDigital Libraries

Summary

The authors studied how open large language model (LLM) research shares various related resources like smaller model versions and code across many platforms. They found it is hard for digital libraries to properly identify and preserve these interconnected resources as complete scholarly objects. By creating a framework to measure how well these resources can be curated, they analyzed thousands of public repositories and papers, discovering that few contain enough linked information to be fully preserved and cited. The authors suggest a minimal set of information fields and roles for different organizations to improve how LLM artifacts are managed and referenced.

Large Language ModelsModel AdaptersQuantized CheckpointsDigital LibrariesScholarly RecordsModel HubsCuratabilityScholarly LinkagePreservationBibliographic Control

Authors

Yiyi Lu, Yilai Qian, Yucheng Jin

Abstract

Open Large Language Model (LLM) research increasingly produces compact and derived artifacts, such as adapters, quantized checkpoints, merged models, and distilled variants, that are distributed across papers, model hubs, model cards, code repositories, and release statements. Although these artifacts are publicly visible, digital libraries often lack sufficient evidence to identify, preserve, and cite them as coherent scholarly objects. We introduce a framework that conceptualizes curatability as a record-level property of distributed scholarly records and operationalizes it through four evidence dimensions: artifact identity, scholarly linkage, upstream evidence, and release assets. Guided by this framework, we conduct the first collection-scale characterization of open LLM curatability using a May 2026 snapshot of 191,375 public Hugging Face repositories and a core corpus of 2,214 scholarly papers. Our results reveal a pronounced visibility-to-curatability funnel. While 90.7% of paper records contain at least one useful curation signal, only 18.1% combine usable upstream evidence with concrete release evidence, and only 6.1% provide sufficiently coordinated evidence to support high-curatability records. Based on these findings, we derive a minimal seven-field curatable record and complementary responsibilities for model hubs, scholarly indexes, and digital libraries, providing practical guidance for improving the preservation and bibliographic control of open LLM artifacts.