Papers for

software platform engineers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Tool call safety improved by matching calls to implementation authority

When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents

Abstract: Tool-enabled agents form calls from model-visible interfaces, while hosts later select their implementation. Standard dispatch omits the descriptor-handler relation. An unchanged and schema-valid call can therefore acquire a different security effect during rollout, reconnect, or delayed approval. We call this failure schema-epoch drift. We present formation-consistent dispatch (FCD), which connects implementation analysis to execution authority. Reviewed profiles produce provenance-bound over-approximations of declared in-scope effects from official source. Under a closed-target approval policy, a verifier applies each formed call to a summary and captures a successor only when its effects fit the call's security contract. Atomic admission and a final-hop fence preserve this decision to the effect. The exact source retains priority, and the captured successor becomes eligible only after source retirement. Stock releases and deployment changes reproduced the failure. Four profiles covered 32 official releases: 29 required no release-specific change and three escalated. A frozen 16-release expansion matched a separate source oracle. In a preregistered stock comparison, FCD completed all three pending calls whose effect remained private and blocked all three whose omission became public. Exact pinning and release-wide denial stopped all six calls, while release-wide approval completed all six but produced three public effects. A separate lifecycle experiment carried a formation-captured certificate across source retirement. The same safe certificate installed later governed new formations without expanding the pending call's authority.

Mon 28 SeptArtificial IntelligenceCryptography and SecuritySoftware Engineering
The gist
Sometimes, software tools call other parts of a program in ways that look the same but actually do different things, which can cause problems with security and trust. The authors show this happens when the connection between the call and its handler changes over time, called schema-epoch drift. They propose a method called formation-consistent dispatch (FCD) that keeps track of how calls relate to their actual implementations and the security permissions involved, so calls only proceed if they meet safety checks. Their tests across many software releases showed FCD can correctly allow safe calls and block unsafe ones, even as software changes over time.
Open → 2609.35088v1

AI agent performance varies widely with resource use and task type

Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics

Abstract: LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers. This execution model can make the optimization of agent serving difficult because latency, local resource demand, and container bottlenecks inter-mix across requests. However, the current agent ecosystem runs without much consideration of resource dynamics, which results in significant waste of the precious resources. This paper analyzes the resource inter-mix of AI agents for three representative tasks: retrieval-augmented question answering, web search, and software coding. To this end, we characterize the latency with respect to the resource dynamics of processing multiple requests and tasks concurrently. Our measurements show that agents have a wide range of behaviors depending on tasks, so that even the same tool can differ substantially in resource dynamics. We also find that running multiple requests concurrently exposes task-dependent bottlenecks in resource dynamics such as CPU, disk I/O, and memory. Furthermore, we uncover that faster LLM responses or more CPU cores do not always accelerate agents. Based on these observations, we demonstrate new optimization opportunities that exploit the resource dynamics of tasks: CPU-aware tool admission and task-aware CPU allocation. Our results show that the latency of CPU-sensitive agent tasks improves $\sim$5.4$\times$, and the average latency across multiple tasks is reduced $\sim$32% compared to native agents.

Thu 17 SeptArtificial IntelligencePerformance
The gist
AI agents that answer questions or write code use a mix of local tools and remote language models. The authors found that these agents behave very differently depending on the task, and running many requests at once can cause bottlenecks with the computer’s CPU, memory, or disk. Surprisingly, faster language model responses or adding more CPU power doesn't always speed things up. By understanding these resource patterns, the authors showed it is possible to manage computing resources better and reduce waiting times significantly.
Open → 2609.19947v1

Netkit improves container network speed by removing delays

Netkit: Specializing Linux Packet Delivery for Container Networks

Abstract: Cloud-native microservices architectures rely on network namespaces for isolation, with the overhead of container communications remaining a critical performance bottleneck. While colocating containers on the same host mitigates some of this overhead, it cannot match the performance of communication within a single network namespace. Existing solutions either require application rewrites or fail to support the full Linux network stack expected by containerized applications. In this paper, we present netkit, an eBPF-based datapath that specializes the Linux networking stack to eliminate redundant backlog queue traversals during network namespace transitions. netkit leverages eBPF to transparently redirect packets between namespaces, bypassing unnecessary buffering while preserving compatibility with existing container applications. Our implementation in the Linux kernel, integrated with minimal changes to the Cilium network plugin for Kubernetes, improves throughput by up to 37\% and achieves parity between container-to-container and process-to-process communications, effectively closing the performance gap introduced by namespace isolation.

Wed 16 SeptOperating SystemsNetworking and Internet Architecture
The gist
Communicating between small programs called containers usually slows down because of how data is handled across separate network spaces. The authors created netkit, which uses a special Linux feature called eBPF to let data move directly between containers without waiting in unnecessary queues. This change lets containers talk to each other as fast as programs running together in the same space, improving speed by up to 37%. netkit works without changing the container programs themselves and fits easily into existing cloud systems.
Open → 2609.18633v1