Papers for
software platform engineers
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Tool call safety improved by matching calls to implementation authority
When Valid Tool Calls Change Meaning: Formation-Consistent Dispatch for LLM Agents
Abstract: Tool-enabled agents form calls from model-visible interfaces, while hosts later select their implementation. Standard dispatch omits the descriptor-handler relation. An unchanged and schema-valid call can therefore acquire a different security effect during rollout, reconnect, or delayed approval. We call this failure schema-epoch drift. We present formation-consistent dispatch (FCD), which connects implementation analysis to execution authority. Reviewed profiles produce provenance-bound over-approximations of declared in-scope effects from official source. Under a closed-target approval policy, a verifier applies each formed call to a summary and captures a successor only when its effects fit the call's security contract. Atomic admission and a final-hop fence preserve this decision to the effect. The exact source retains priority, and the captured successor becomes eligible only after source retirement. Stock releases and deployment changes reproduced the failure. Four profiles covered 32 official releases: 29 required no release-specific change and three escalated. A frozen 16-release expansion matched a separate source oracle. In a preregistered stock comparison, FCD completed all three pending calls whose effect remained private and blocked all three whose omission became public. Exact pinning and release-wide denial stopped all six calls, while release-wide approval completed all six but produced three public effects. A separate lifecycle experiment carried a formation-captured certificate across source retirement. The same safe certificate installed later governed new formations without expanding the pending call's authority.
AI agent performance varies widely with resource use and task type
Not All AI Agents Are Equal: Characterizing Resource and Performance Dynamics
Abstract: LLM-based AI agents process user requests through iterative reasoning and tool execution, often involving the invocation of remote LLM APIs with local tool containers. This execution model can make the optimization of agent serving difficult because latency, local resource demand, and container bottlenecks inter-mix across requests. However, the current agent ecosystem runs without much consideration of resource dynamics, which results in significant waste of the precious resources. This paper analyzes the resource inter-mix of AI agents for three representative tasks: retrieval-augmented question answering, web search, and software coding. To this end, we characterize the latency with respect to the resource dynamics of processing multiple requests and tasks concurrently. Our measurements show that agents have a wide range of behaviors depending on tasks, so that even the same tool can differ substantially in resource dynamics. We also find that running multiple requests concurrently exposes task-dependent bottlenecks in resource dynamics such as CPU, disk I/O, and memory. Furthermore, we uncover that faster LLM responses or more CPU cores do not always accelerate agents. Based on these observations, we demonstrate new optimization opportunities that exploit the resource dynamics of tasks: CPU-aware tool admission and task-aware CPU allocation. Our results show that the latency of CPU-sensitive agent tasks improves $\sim$5.4$\times$, and the average latency across multiple tasks is reduced $\sim$32% compared to native agents.
Netkit improves container network speed by removing delays
Netkit: Specializing Linux Packet Delivery for Container Networks
Abstract: Cloud-native microservices architectures rely on network namespaces for isolation, with the overhead of container communications remaining a critical performance bottleneck. While colocating containers on the same host mitigates some of this overhead, it cannot match the performance of communication within a single network namespace. Existing solutions either require application rewrites or fail to support the full Linux network stack expected by containerized applications. In this paper, we present netkit, an eBPF-based datapath that specializes the Linux networking stack to eliminate redundant backlog queue traversals during network namespace transitions. netkit leverages eBPF to transparently redirect packets between namespaces, bypassing unnecessary buffering while preserving compatibility with existing container applications. Our implementation in the Linux kernel, integrated with minimal changes to the Cilium network plugin for Kubernetes, improves throughput by up to 37\% and achieves parity between container-to-container and process-to-process communications, effectively closing the performance gap introduced by namespace isolation.