Papers for

software platform developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Client resolved generation protects data privacy in cloud language models

Learning to Refer: Client-Resolved Generation for Privacy-Aware Language Models

Abstract: Cloud-based large language models (LLMs) require users to disclose plaintext data to service providers, creating privacy risks in sensitive domains. Existing privacy-preserving approaches often trade utility for protection, incur substantial computational or communication overhead, remain vulnerable to reconstruction from intermediate representations, or protect only a subset of the training and inference pipeline. We introduce Client-Resolved Generation (CRG), a genera- tion interface that separates server-side generation from the lexical realization of input-derived content. The client transmits only pooled and noise-perturbed rep- resentations, while input-derived output content is represented using request-local positional references and resolved to its original strings only on the client. This interface protects private input and input-derived output content during both train- ing and inference while allowing the service provider to keep its proprietary model parameters hidden from the client. At the same time, exact lexical reuse remains possible without directly exposing the reused content on the provider-visible gen- eration path. We evaluate CRG on medical and document-grounded QA, sensi- tive identifier transfer, and tool calling, together with reconstruction and raw-logit leakage analyses. On SealTools, CRG improves complete-call exact match from 57.3% to 79.9% over the input-privacy framework PPFT, with larger gains as more required output content can be resolved through references. Together, these results show that CRG provides a practical interface for privacy-sensitive cloud LLMs by reducing plaintext exposure across both input and output pathways while preserv- ing task utility and server-side model confidentiality.

Sat 26 SeptCryptography and SecurityArtificial Intelligence
The gist
Cloud services using large language models often require users to send unprotected text, risking privacy. The authors propose Client-Resolved Generation (CRG), where sensitive input and output content is encoded and referenced without exposing raw text to the server. This approach keeps user data private during both training and use, while also safeguarding the service provider’s model. Tests show CRG improves accuracy and reduces privacy risks without hurting the model’s usefulness.
Open → 2609.32706v1

Provenance based runtime guard stops cascading attacks on llm agents

AGATE: Provenance-Based Runtime Defense Against Compositional Attacks on LLM Agents

Abstract: LLM agents can produce harmful effects through sequences of ordinary operations. Judging such actions requires establishing both the authority that permits them and the origin of the data they carry. We present AGATE, an authorization and data-provenance gate at instrumented agent-harness boundaries. Operator declarations and host approval events ground authorization; delegated actions are constrained by grants that bind to exact parameters, expire, and permit a limited number of uses. Source registration connects observed inputs to subsequent transfers, while an effect ledger tracks repeated requests. Deterministic checks make decisions without an LLM in the decision path and retain their grounds with execution evidence for forensic replay. Adapters integrate three production harnesses -- DeepSeek Harness, OpenCode, and OpenClaw -- without modifying host code, translating each host's native observation and veto points into a single shared gate interface; the judgment core is identical in all three, and only enforcement depth differs. Our evaluation combines 153 exercised attack-chain records with deployment, utility, and reconstruction experiments. The deployment observations expose how tool declarations and data checks govern business actions, including a bypass through parameter rewriting. Six of eleven benign file-processing scenarios contain denial events, revealing the utility cost of content-based provenance policies. Across 252 runs on 63 sanitized scenarios, replay agrees with live graph projections for all 63 scenarios on each of two platforms. These results establish the feasibility of provenance-based runtime judgment and identify content transformation, legitimate reuse, and observation coverage as concrete limits.

Fri 25 SeptCryptography and SecuritySoftware Engineering
The gist
Large language model (LLM) agents can cause harm by chaining normal actions that look safe on their own. The authors created AGATE, a tool that checks who allowed each action and where the data comes from to decide if the action is okay. It watches agent boundaries and records detailed evidence so decisions can be replayed for review. AGATE works with three existing LLM frameworks without changing their code and was tested on many attack and normal scenarios to verify its accuracy and limits.
Open → 2609.30830v1

Linking literature and infrastructure improves research data reuse

A framework for linking literature-based knowledge integration and infrastructure-supported knowledge integration: Opportunities and challenges from a case study

Abstract: Integrating knowledge across disciplines is central to sustainability research, yet most evidence-synthesis methods rely on findings as reported in publications, limiting verification and reuse of underlying data and workflows. We develop a conceptual framework linking literature-based and infrastructure-supported knowledge integration, using a systematic review case study to examine when integration can extend beyond reported findings. We reviewed 37 studies on climate change, violent conflict, and household food security. Literature-based synthesis enabled integration across all included studies, whereas access to reusable outputs was limited: over half provided no data availability statement, 27% reported availability upon request, but reusable data and workflows were available for only 8%. To explore infrastructure-supported integration, we used the TIB Knowledge Loom to represent studies with accessible data and code as machine-readable outputs, and produced a knowledge gap map (KGM) from manually extracted and Loom-derived data, comparing manual and infrastructure-supported synthesis. Where outputs were reusable, synthesis could be produced directly from data and workflows rather than from publications. These findings show that literature-based synthesis can be complemented by infrastructure-supported integration where outputs are accessible and usable, and that advancing knowledge integration depends not only on infrastructures but on making data, code, and workflows accessible, executable, and reusable.

Thu 24 SeptDigital LibrariesDatabases
The gist
Combining knowledge from many studies helps solve big problems like climate change, but most research only shares final results, not the data or steps behind them. The authors looked at 37 studies on climate, conflict, and food security to see how often researchers share data and code openly. They found very few studies make reusable data and workflows available, which limits building on past work. Using a tool called TIB Knowledge Loom, they showed better knowledge integration is possible when data and methods are accessible and machine-readable.
Open → 2609.29161v1