Papers for

high-performance computing teams

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Loome enables fast tile program scheduling without tuning on dataflow chips

Schedules Are Solvable Symbols: Tuning-Free Compilation of Tile Programs on Dataflow Architectures

Abstract: Modern AI and HPC accelerators increasingly expose dataflow features: software-visible mechanisms for data movement and overlap, such as inter-core communication through the on-chip network and intra-core asynchronous pipelining. These features shift scheduling responsibility from hardware to the compiler, and because placement, movement, and synchronization become software-visible, they also make the performance of static schedules predictable. Yet high performance on such hardware still relies on vendor-engineered kernel libraries or profile-based auto-tuning, whose embedded expert knowledge transfers poorly across architectures and algorithms. We present Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures. The central idea is to treat tile-based SPMD compilation as a hardware-explicit static optimization problem. Loom enumerates discrete spatial-mapping and communication candidates while keeping value parameters, such as tiling factors and pipeline knobs, symbolic within each candidate. From an explicit hardware description, it derives symbolic legality constraints and latency expressions, formulates one CP-SAT problem per schedule candidate, and jointly solves inter-core dataflow, intra-core asynchronous scheduling, and block sizes at compile time. On two Tenstorrent generations, Wormhole and Blackhole, Loom matches or exceeds the vendor-optimized TTNN library on GEMM, Flash Attention, and Flash Decode, out of the box and without per-shape profiling or profile-based platform-specific schedule tuning. These results suggest that hardware-derived symbolic compilation provides a retargetable alternative to profiling-based tuning for spatial dataflow architectures while remaining interpretable by keeping optimization decisions traceable to source-level symbols.

Thu 24 SeptProgramming Languages
The gist
Many specialized computer chips need careful planning to run programs fast, relying on expert-made libraries or slow testing to find the best settings. The authors introduce Loom, a new way to automatically create these plans by imagining all valid options and solving equations about timing and communication without trial runs. Loom works right away on advanced chips to run important tasks efficiently without extra tuning. This method can save time and adapt better to different hardware designs.
Open → 2609.29219v1

TileBench compares tile-based AI programming models for GPU speed

TileBench: A Controlled Benchmark for Performance Evaluation and Bottleneck Diagnosis of Tile-Based Programming Models

Abstract: Tile-based programming models, such as Triton and cuTile, aim to simplify high-performance kernel development, but their practical performance, tuning behavior, and usability remain difficult to compare systematically. We present TileBench, a controlled benchmark for evaluating Triton and cuTile on NVIDIA B200 GPUs under matched operator semantics and comparable implementation structures. TileBench contains 45 operators covering diverse AI-kernel patterns and memory/computation behaviors. Each task provides a PyTorch reference, verified Triton and cuTile implementations, standardized data-types (dtype) and input-size sweeps, default and autotuned configurations, roofline-based metrics, and profiling-guided diagnosis. Our evaluation shows that performance gaps are workload-dependent: cuTile excels on a small cluster of Tensor-Core/TMA-friendly kernels, while Triton is stronger on many irregular, streaming, and bandwidth-bound operators. We further evaluate LLM-generated cuTile and Triton kernels and find that Triton is consistently more token-efficient than cuTile under the same iterative refinement protocol. TileBench is publicly available at https://github.com/Deep-Learning-Profiling-Tools/Tilebench.

Thu 24 SeptPerformanceProgramming Languages
The gist
Building fast programs for GPUs is tricky because different programming models behave differently depending on the task. The authors created TileBench, a test set of 45 AI-related tasks with matching implementations in two popular tile-based programming models, Triton and cuTile. They measured how each model performed on NVIDIA B200 GPUs and found that cuTile works better for some specialized tasks, while Triton performs better on many others, especially those with irregular or bandwidth-heavy operations. They also tested kernels generated by language models and saw Triton was more efficient in refining code. TileBench offers a fair way to compare these models for anyone developing GPU programs.
Open → 2609.29067v1

Tetris improves scheduling for photonic switch networks in computing

Tetris: Circuit Scheduling for Rearrangeably Non-Blocking Photonic Interconnects

Abstract: Reconfigurable photonic interconnects are emerging as a promising communication architecture for next-generation distributed computing. Yet, most circuit schedulers are designed around an idealized view of the interconnect as either blocking or strictly non-blocking. Practical scalable designs are often rearrangeably non-blocking (RNB), with connections routed through networks of internal $2\times2$ switches. This changes the scheduling problem fundamentally: establishing a new connection can force existing connections to be rerouted, trigger state changes across multiple internal switches, and impose reconfiguration delay on otherwise unrelated traffic. We present Tetris, a circuit scheduling algorithm for RNB photonic interconnects. Tetris builds on two observations. First, All-to-All demands are often not doubly stochastic, leaving a small number of endpoints as communication bottlenecks. Second, reconfiguration delay can be large enough to change which connection should be scheduled next. Tetris prioritizes bottleneck endpoints using their remaining communication and reconfiguration work, while selecting and routing matchings to preserve ongoing connections whenever possible. Matchings ensure progress, but connections are scheduled independently, allowing completed connections to be replaced without matching-wide barriers and incurring delay only at switches whose states change. Our simulation and hardware-emulation results show that Tetris reduces All-to-All demand completion time by up to $6.6$x over Birkhoff--von Neumann-based scheduling and by $30$% over Sunflow. More broadly, RNB interconnects raise new questions in multi-tenant scheduling and routing for partial reconfiguration, which we discuss at the end of the paper.

Mon 21 SeptNetworking and Internet Architecture
The gist
Photonic interconnects use light to send data between computers, but managing connections is tricky because new connections can disrupt existing ones. The authors developed Tetris, a clever scheduling method that prioritizes busy parts of the network and tries to avoid unnecessary changes during data transfer. Their approach speeds up communication by up to 6.6 times compared to older methods, especially for all-to-all communication tasks. This work highlights new challenges and solutions for networks that can rearrange their connections dynamically.
Open → 2609.25434v1

Ai systems improve managing scientific computing workflows and reproducibility

AI-Driven Scientific Computing Workflows: A Systems Review of Orchestration, Execution, Reproducibility and Provenance

Abstract: Artificial intelligence (AI) is increasingly embedded within scientific computing workflows that combine simulation, data processing, optimisation, visualisation and experimental or observational components. Learned models may serve as explicit workflow components, retain persistent state and, in adaptive settings, influence subsequent computation. Existing work has characterised scientific workflow management systems, dynamic and steered workflows, AI--HPC coupling motifs and the machine-learning lifecycle, although these areas are often discussed separately. This review brings them together from a systems perspective. We distinguish conventional scientific workflows, machine-learning pipelines, AI-coupled high-performance computing (HPC) workflows and broader automated research workflows, and propose a continuum describing the depth of AI participation from a computational stage to co-adaptive workflow control. The associated systems requirements are organised around five concerns: control and orchestration; compute and execution; data and model state; reproducibility and provenance; and governance and assurance. Representative systems and applications include AI-steered molecular simulation, drug and materials discovery, simulation--surrogate coupling and distributed self-driving laboratories. Workflow-level evaluation is considered in terms of scientific progress, execution cost, data movement, resource use, resilience and decision traceability. We conclude by identifying open problems in dynamic workflow representation, state-aware recovery, heterogeneous scheduling, interoperable data planes, model-mediated decision provenance and reproducible adaptive execution.

Fri 18 SeptDistributed, Parallel, and Cluster Computing
The gist
Scientific computing combines many steps like simulations, data analysis, and experiments. The authors reviewed how artificial intelligence (AI) helps manage these complex workflows, making them more adaptive and efficient. They organized systems needs around control, execution, data handling, reproducibility, and trust. The paper highlights uses like drug discovery and automated labs, and points out future challenges in making AI-driven workflows more reliable and transparent.
Open → 2609.21162v1

Simd vectorization triples performance in generating permutations efficiently

Parallelizing the Factorial Space: 3x SIMD Acceleration of the Steinhaus-Johnson-Trotter Algorithm via Dual-Lane AVX2 Execution

Abstract: This paper presents a high-performance SIMD acceleration framework for the Steinhaus-Johnson-Trotter permutation generation algorithm, targeted at modern x86-64 architectures using the AVX2 instruction set. By exploiting a novel combinatorial space partitioning with pre-calculated index offsets combined with single-cycle vector byte shuffling (_mm256_shuffle_epi8), our dual-lane vectorized implementation processes two independent, concurrent permutation streams within a single 256-bit YMM register under a uniform execution mask. Empirical evaluations demonstrate a 3$x$ performance throughput increase over both Donald Knuth's Algorithm P (TAOCP Vol 4A), which we previously accelerated by 3$x$ in scalar code, and the recent Ring-Cascade algorithm by Yusheng Hu. The proposed software architecture maintains cross-compiler compliance, completely avoids store-forwarding memory stalls during hot loops, and is validated up to order $n=11$ with a benchmark performance of ~1.27 billion CPU cycles for $n=13$ on native hardware.

Mon 7 SeptData Structures and Algorithms
The gist
Generating all possible orders (permutations) of items is a common computing challenge that can be very slow for larger lists. The authors developed a new way to speed up a well-known algorithm by running two streams of computations in parallel using advanced CPU instructions called AVX2. This approach shuffles and manages data inside the processor's registers to process permutations much faster without slowing down memory access. Their method runs about three times faster than previous best-known techniques for sequences up to size 13 on typical modern processors.
Open → 2609.07862v1

Gpu acceleration cuts latency and energy for neuron core mapping

GPU-Accelerated Hypergraph Partitioning and Placement to Map SNNs on Neuromorphic Hardware

Abstract: SNNs running on neuromorphic hardware use spikes to achieve sparse and energy-efficient communication over a mesh of cores. In turn, system performance heavily depends on the assignment of neurons to cores: the mapping. Since hardware features inter-core multicast and intra-core replication of spikes, we model SNNs as hypergraphs to exploit both opportunities for reducing communication traffic. Mapping thus comprises two NP-hard problems: hypergraph partitioning and placement on the lattice of cores. High-quality solutions to both are critical, yet increasingly difficult as networks scale to millions of neurons. Therefore, we propose a GPU-accelerated pipeline for SNN mapping: a multi-level partitioning scheme is devised around hardware constraints, while placement is initialized through recursive bisection, followed by refinement pulling together strongly connected cores through repeated swaps. Model-based experiments show upwards of 16% lower latency and 42% lower energy for spike movements over existing sequential tools, while our parallel mapper is on average 18-280x faster.

Mon 7 SeptDistributed, Parallel, and Cluster Computing
The gist
Spiking neural networks use bursts of activity to communicate efficiently across many small processors. How these neurons are assigned to the processors affects how fast and energy-efficient the system runs. The authors created a way to speed up this assignment on graphics processors by breaking the problem into pieces and improving connections between processors. Their approach reduces the time data spends moving around and uses less energy compared to older methods, while running much faster.
Open → 2609.07577v1