Papers for

iot platform developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

AgentWare automates deploying AI agents across edge to cloud networks

AgentWare: Automating the Lifecycle of Agentic Applications across the Edge-to-Cloud Continuum

Abstract: Deploying LLM-enabled agentic applications across the Edge-to-Cloud continuum remains challenging due to hardware heterogeneity, deployment complexity, limited observability, and the lack of systematic evaluation methods. Existing solutions address agent development, observability, or benchmarking separately, offering limited support for the full lifecycle of distributed agentic applications. This paper presents AgentWare, an AgenticOps framework that automates the provisioning, deployment, observability, and evaluation of agentic applications across Edge-to-Cloud infrastructures. AgentWare introduces an end-to-end lifecycle pipeline that automatically prepares heterogeneous execution environments, transforms user-defined agent implementations into distributed applications, deploys agent components across the continuum, and performs unified collection of execution traces, infrastructure telemetry, and evaluation metrics. The framework further supports automated semantic evaluation through LLM-as-a-Judge workflows and generates reproducible reports covering correctness, performance, resource utilization, and energy consumption. We demonstrate the applicability of AgentWare through a distributed book assistant agent deployed across real Edge-to-Cloud infrastructure under multiple deployment and model configurations. The results show that AgentWare enables systematic experimentation and evaluation of distributed agentic applications while significantly reducing the manual effort required for deployment, instrumentation, and analysis.

Mon 28 SeptDistributed, Parallel, and Cluster ComputingMachine Learning
The gist
Running AI programs that work together on different computers from small devices to big servers is hard because these machines are very different and tricky to manage. The authors created AgentWare, a tool that automatically sets up, runs, watches, and tests these AI programs across many kinds of machines. It helps collect data about how well the AI agents work and how much energy or resources they use. They showed how AgentWare makes it easier to experiment and learn from these systems by testing a book helper AI spread across edge and cloud computers.
Open → 2609.34586v1

Internet of agents gets scalable trust system for identity and discovery

A Scalable Trust Discovery Architecture for the Internet of Agents

Abstract: The Internet of Agents is expected to enable large numbers of autonomous agents to discover, verify, and collaborate with each other across heterogeneous platforms. However, current agent protocols mainly address tool invocation and inter-agent communication, leaving scalable agent registration, trustworthy identification, and capability-oriented discovery largely unresolved. To address this, this paper proposes a scalable trust discovery architecture for the Internet of Agents. The proposed architecture adopts a hierarchical and distributed design consisting of three layers: Agent Root for trusted registry governance, Agent Registry for agent registration and metadata publication, and Agent Resolver for distributed capability discovery and trust-aware resolution. The architecture further introduces a registry-suffix-anchored composite identity scheme, which binds an agent native identifier to a trusted registry suffix to generate a globally discoverable identity. It also incorporates a dual-certificate and multi-level authentication mechanism to strengthen identity trust among agents. We implement a prototype and evaluate it through large-scale agent registration and resolution experiments. The prototype achieves an average registration latency of 58ms and an average discovery latency of 25ms, and it supports more than 19,000 registration requests per second and more than 29,000 agent discovery requests per second. These results demonstrate the feasibility of the proposed architecture, providing a practical approach toward scalable and identity-trusted agent ecosystems in the Internet of Agents.

Thu 17 SeptCryptography and SecurityArtificial Intelligence
The gist
Many autonomous agents on the internet need to find and trust each other to work together well. Existing systems mostly focus on communication but don’t handle registering agents or verifying their identities at a large scale. The authors built a new architecture that organizes agents in layers and links each agent’s ID to a trusted registry, which helps everyone find and trust each other efficiently. Their prototype handled tens of thousands of registrations and discoveries per second with very low delays, showing this approach can work for huge numbers of agents.
Open → 2609.20095v1

Hierarchical system speeds up predictions on streaming data with edge and cloud

Streaming Hierarchical Inference with Tabular Foundation Models

Abstract: Tabular Foundation Models (TFMs) have recently demonstrated strong predictive performance through in-context learning, but their deployment in high-throughput data streams remains challenging due to communication overhead and latency. We propose \textit{HINT}, a hierarchical inference framework that combines edge-based retrieval with cloud-based TFM inference. A graph-based approximate nearest neighbor memory maintained over a sliding window provides local predictions and uncertainty estimates, allowing confident samples to be processed locally while uncertain instances are selectively offloaded, together with their retrieved context, to a cloud-hosted TFM. The framework exposes an offloading threshold and a neighborhood retrieval policy that can be varied to balance predictive performance and communication cost. Experiments show \textit{HINT} consistently identifies favorable trade-offs.

Mon 7 SeptMachine Learning
The gist
Fast and accurate predictions on data streams are hard when sending everything to the cloud causes delays and extra communication. The authors propose a method called HINT that first tries to make predictions on a local device using a recent memory of past data. When the local system is unsure, it sends just those uncertain cases along with related information to a powerful cloud model for better analysis. Their approach lets users choose how much to rely on the local or cloud parts, balancing speed and accuracy. Experiments show this method finds a good middle ground between quick responses and prediction quality.
Open → 2609.07956v1