Papers for
platform maintainers
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Most apps update large language models after failures despite advance notices
When the Model Retires: An Empirical Study of LLM Migration in Open-Source Applications
Abstract: Applications built on commercial large language model (LLM) APIs depend on model versions that providers retire on their own schedule, with notice periods ranging from one year to two weeks. We ask what actually happens to applications when a model is retired. We mine GitHub for commits that migrate away from officially deprecated models and endpoints of OpenAI, Anthropic, and Google, matching each commit to the provider's published announcement and shutdown dates. From 22,555 commits in 17,703 non-fork repositories (2024-2026), 5,139 are matched to an official event; two independent coders validated a stratified sample of 300 (kappa = 0.89-0.95), and we reweight all estimates by their labels. We find that an estimated 82% (95% CI 79-84) of migrations away from retired models were committed after the shutdown date - after the application had started failing - regardless of repository popularity, prior retirement experience, or the presence of a provider-abstraction layer. The share tracks the provider's notice policy: 89% for Anthropic's 60-114-day notices versus 13% for OpenAI's one-year Assistants API notice, and each e-fold increase in notice length reduces the odds of post-shutdown migration by about three quarters. Model identifiers are hard-coded in 94% of migrating applications, migration effort scales from a median of 6 added lines for prompt-only applications to nearly 700 for fine-tuned ones, and only 8% of migrations switch provider. We release the dataset and pipeline and discuss implications for deprecation policy, dependency-risk assessment of LLM products, and tooling.
Agent plugins rarely follow standards and still clash with others
Packaged, But Not Portable: Why Conforming to the Agent Plugin Standard Is Rare, and Why Conforming Would Not Be Enough
Abstract: Coding agents are extended by plugins: installable bundles that ship skills, sub-agents, commands, hooks, and tool servers. On 24 July 2026, an open specification (Agent Plugins v1.0.0) standardised how such a bundle is laid out and described, so that one plugin could run on any agent. We ask the two questions a practitioner would ask of it: is the ecosystem adopting the standard, and if a plugin did conform, would that be enough to make it work alongside the other plugins a user has installed? We answer both by building AgentPluginZoo, a provenance-tracked corpus of 68,072 plugin bundles across 30,655 repositories, released with its discovery ledger, scoring code, and analysis. Only 6.2% validate, but the gap is shallow rather than structural: 96.6% would load after adding one missing boilerplate field. The real cost lands elsewhere. 40.2% would load while the specification obliges the client to discard fields their authors wrote, mostly declarations of what the plugin ships. Moreover, conformance settles nothing for the second question: 81% of capability-exporting bundles share a name with another plugin, with no namespace or precedence rule to decide which one answers. This paper argues the community standardised a packaging format when composition needs a model, names the four concepts such a model must add - qualified capability identity, a declared capability surface, a precedence rule, and inter-plugin relations - and shows they fit an additive v1.1 profile of the same specification rather than a competing standard. Recommendations follow for practitioners packaging extensions today and for the people evolving the standard, chief among them that conformance must be made observable before it can become common. The corpus and code are available at https://github.com/tezansahu/agentpluginzoo