Papers for

data center managers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Trace improves cable tracking accuracy amid clutter and crossings

TRACE: Interactive Bi-Directional Tracing of Monochrome Cables Amid Clutter

Abstract: Accurate state estimation (tracing) of Deformable Linear Objects (DLOs) such as cables is a critical challenge for data centers, manufacturing, construction, homes, and surgery, where precise cable management directly impacts operational safety and efficiency. However, resolving the state of multiple monochrome cables amid foreground and background clutter poses challenges due to occlusions, overlap, and ambiguous crossings. We present Two-way Routing And Cable Estimation (TRACE), which combines bi-directional cable tracing with interactive perception primitives-Divergence Push and Cluster Dilation-to actively resolve ambiguities. Evaluation with 110 physical experiments suggests that TRACE can increase the percentage of cable length correctly traced in complex scenarios (with up to 4 cables and 40 crossings) from ~60% with the strongest prior method, HANDLOOM 2.0, to ~90%, outperforming RT-DLO, Nano Banana Pro, and ChatGPT 5.2 as well. For a trial run on a workstation with an NVIDIA GeForce RTX 4090 GPU, the average computation time is 0.4 seconds per cable. Project website: https://trace-paper.github.io/.

Thu 24 SeptRobotics
The gist
Tracing cables that overlap or cross each other is tricky, especially when they look the same color and get tangled. The authors created TRACE, a method that traces cables in two directions and uses interactive steps to clear up confusion. This approach tracks cables more accurately than previous methods, even in difficult setups with many cables and crossings. It works fast, taking less than half a second per cable on a powerful computer.
Open → 2609.29103v1

Interactive 3d digital twin helps visualize supercomputer hardware and activity

Object Model Analysis of a Supercomputer with Digital Twin

Abstract: Operators and developers need a mental model of both the structure and the live behavior of a large supercomputer, but its physical layout, logical organization, and streams of per-node telemetry are difficult to relate to one another, making it hard to trace a metric or event back to a specific hardware component. We present DAT, an interactive three-dimensional digital analytics twin of a compute cluster built in a real-time game engine, Unreal Engine. DAT expands a compact, parametric description of a supercomputer, a reusable Digital Twin Prototype (DTP), into a navigable Digital Twin Instance (DTI) that mirrors its physical containment hierarchy of racks, chassis, blades, and network links, encoding each node's role and health in its appearance, while a lightweight event-driven simulator animates job and hardware activity over a virtual clock. Our current implementation adds a two-path node-selection mechanism, unifying direct 3D pointing with command-shell queries, that opens an in-world visual-analytics panel beside any selected component showing summary statistics and live, time-varying metrics. We describe this architecture, report qualitative behavior from the working prototype, and outline the path toward driving the panels with recorded telemetry and in-situ anomaly detection.

Fri 11 SeptDistributed, Parallel, and Cluster Computing
The gist
Supercomputers are complex machines made of many parts, and it can be hard to understand how all these parts work together or spot problems. The authors created a 3D digital twin, a virtual model, that shows the supercomputer’s physical setup like racks and network links, along with live activity and health information. This model runs in a game engine and lets users explore and select parts either by clicking in 3D or using commands, showing detailed stats and ongoing metrics. This helps operators and developers quickly connect performance or problems to the right hardware piece.
Open → 2609.13571v1

Mathematical model improves scheduling for large scientific workflows

Mathematical Modeling of a Cognitive Continuum Digital Shadow for Large-Scale, Cross-Facility Workflows

Abstract: We present the mathematical foundations of a \emph{Cognitive Continuum Digital Shadow} (CCDS), a decision-support layer between users and the cross-facility infrastructure---instruments, networks, data stores and compute centers---of exascale and post-exascale scientific workflows. The CCDS couples a state-space representation of the continuum with multistage stochastic programming, so that deployment scenarios can be explored and optimized \emph{before} jobs are launched. This allows operators and users to quantify the cost, makespan and energy trade-offs of a workflow under uncertain resource availability, and hedge their decisions accordingly. We formulate the underlying optimization as a multimode, resource-constrained, stochastic supply-chain network design problem and demonstrate it on a realistic genomics workflow scheduled across heterogeneous HPC and data-center resources. This is the first of three papers; the second treats the underlying software architecture and the third reports large-scale use-cases.

Mon 7 SeptPerformance
The gist
Large scientific workflows often run across many computing centers and instruments, making it hard to plan and manage resources efficiently. The authors present a mathematical approach called the Cognitive Continuum Digital Shadow (CCDS) that helps users and operators understand and predict trade-offs like cost, time, and energy before starting jobs. This approach uses advanced optimization to handle uncertain availability of resources, allowing better decisions on where and when to run parts of a workflow. They demonstrated this method on a genome analysis example using diverse computing resources.
Open → 2609.07275v1