Papers for
high-performance computing teams
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Loome enables fast tile program scheduling without tuning on dataflow chips
Schedules Are Solvable Symbols: Tuning-Free Compilation of Tile Programs on Dataflow Architectures
Abstract: Modern AI and HPC accelerators increasingly expose dataflow features: software-visible mechanisms for data movement and overlap, such as inter-core communication through the on-chip network and intra-core asynchronous pipelining. These features shift scheduling responsibility from hardware to the compiler, and because placement, movement, and synchronization become software-visible, they also make the performance of static schedules predictable. Yet high performance on such hardware still relies on vendor-engineered kernel libraries or profile-based auto-tuning, whose embedded expert knowledge transfers poorly across architectures and algorithms. We present Loom, a tuning-free symbolic compiler framework for tile-based SPMD programs on spatial dataflow architectures. The central idea is to treat tile-based SPMD compilation as a hardware-explicit static optimization problem. Loom enumerates discrete spatial-mapping and communication candidates while keeping value parameters, such as tiling factors and pipeline knobs, symbolic within each candidate. From an explicit hardware description, it derives symbolic legality constraints and latency expressions, formulates one CP-SAT problem per schedule candidate, and jointly solves inter-core dataflow, intra-core asynchronous scheduling, and block sizes at compile time. On two Tenstorrent generations, Wormhole and Blackhole, Loom matches or exceeds the vendor-optimized TTNN library on GEMM, Flash Attention, and Flash Decode, out of the box and without per-shape profiling or profile-based platform-specific schedule tuning. These results suggest that hardware-derived symbolic compilation provides a retargetable alternative to profiling-based tuning for spatial dataflow architectures while remaining interpretable by keeping optimization decisions traceable to source-level symbols.
TileBench compares tile-based AI programming models for GPU speed
TileBench: A Controlled Benchmark for Performance Evaluation and Bottleneck Diagnosis of Tile-Based Programming Models
Abstract: Tile-based programming models, such as Triton and cuTile, aim to simplify high-performance kernel development, but their practical performance, tuning behavior, and usability remain difficult to compare systematically. We present TileBench, a controlled benchmark for evaluating Triton and cuTile on NVIDIA B200 GPUs under matched operator semantics and comparable implementation structures. TileBench contains 45 operators covering diverse AI-kernel patterns and memory/computation behaviors. Each task provides a PyTorch reference, verified Triton and cuTile implementations, standardized data-types (dtype) and input-size sweeps, default and autotuned configurations, roofline-based metrics, and profiling-guided diagnosis. Our evaluation shows that performance gaps are workload-dependent: cuTile excels on a small cluster of Tensor-Core/TMA-friendly kernels, while Triton is stronger on many irregular, streaming, and bandwidth-bound operators. We further evaluate LLM-generated cuTile and Triton kernels and find that Triton is consistently more token-efficient than cuTile under the same iterative refinement protocol. TileBench is publicly available at https://github.com/Deep-Learning-Profiling-Tools/Tilebench.
Tetris improves scheduling for photonic switch networks in computing
Tetris: Circuit Scheduling for Rearrangeably Non-Blocking Photonic Interconnects
Abstract: Reconfigurable photonic interconnects are emerging as a promising communication architecture for next-generation distributed computing. Yet, most circuit schedulers are designed around an idealized view of the interconnect as either blocking or strictly non-blocking. Practical scalable designs are often rearrangeably non-blocking (RNB), with connections routed through networks of internal $2\times2$ switches. This changes the scheduling problem fundamentally: establishing a new connection can force existing connections to be rerouted, trigger state changes across multiple internal switches, and impose reconfiguration delay on otherwise unrelated traffic. We present Tetris, a circuit scheduling algorithm for RNB photonic interconnects. Tetris builds on two observations. First, All-to-All demands are often not doubly stochastic, leaving a small number of endpoints as communication bottlenecks. Second, reconfiguration delay can be large enough to change which connection should be scheduled next. Tetris prioritizes bottleneck endpoints using their remaining communication and reconfiguration work, while selecting and routing matchings to preserve ongoing connections whenever possible. Matchings ensure progress, but connections are scheduled independently, allowing completed connections to be replaced without matching-wide barriers and incurring delay only at switches whose states change. Our simulation and hardware-emulation results show that Tetris reduces All-to-All demand completion time by up to $6.6$x over Birkhoff--von Neumann-based scheduling and by $30$% over Sunflow. More broadly, RNB interconnects raise new questions in multi-tenant scheduling and routing for partial reconfiguration, which we discuss at the end of the paper.
Ai systems improve managing scientific computing workflows and reproducibility
AI-Driven Scientific Computing Workflows: A Systems Review of Orchestration, Execution, Reproducibility and Provenance
Abstract: Artificial intelligence (AI) is increasingly embedded within scientific computing workflows that combine simulation, data processing, optimisation, visualisation and experimental or observational components. Learned models may serve as explicit workflow components, retain persistent state and, in adaptive settings, influence subsequent computation. Existing work has characterised scientific workflow management systems, dynamic and steered workflows, AI--HPC coupling motifs and the machine-learning lifecycle, although these areas are often discussed separately. This review brings them together from a systems perspective. We distinguish conventional scientific workflows, machine-learning pipelines, AI-coupled high-performance computing (HPC) workflows and broader automated research workflows, and propose a continuum describing the depth of AI participation from a computational stage to co-adaptive workflow control. The associated systems requirements are organised around five concerns: control and orchestration; compute and execution; data and model state; reproducibility and provenance; and governance and assurance. Representative systems and applications include AI-steered molecular simulation, drug and materials discovery, simulation--surrogate coupling and distributed self-driving laboratories. Workflow-level evaluation is considered in terms of scientific progress, execution cost, data movement, resource use, resilience and decision traceability. We conclude by identifying open problems in dynamic workflow representation, state-aware recovery, heterogeneous scheduling, interoperable data planes, model-mediated decision provenance and reproducible adaptive execution.
Simd vectorization triples performance in generating permutations efficiently
Parallelizing the Factorial Space: 3x SIMD Acceleration of the Steinhaus-Johnson-Trotter Algorithm via Dual-Lane AVX2 Execution
Abstract: This paper presents a high-performance SIMD acceleration framework for the Steinhaus-Johnson-Trotter permutation generation algorithm, targeted at modern x86-64 architectures using the AVX2 instruction set. By exploiting a novel combinatorial space partitioning with pre-calculated index offsets combined with single-cycle vector byte shuffling (_mm256_shuffle_epi8), our dual-lane vectorized implementation processes two independent, concurrent permutation streams within a single 256-bit YMM register under a uniform execution mask. Empirical evaluations demonstrate a 3$x$ performance throughput increase over both Donald Knuth's Algorithm P (TAOCP Vol 4A), which we previously accelerated by 3$x$ in scalar code, and the recent Ring-Cascade algorithm by Yusheng Hu. The proposed software architecture maintains cross-compiler compliance, completely avoids store-forwarding memory stalls during hot loops, and is validated up to order $n=11$ with a benchmark performance of ~1.27 billion CPU cycles for $n=13$ on native hardware.
Gpu acceleration cuts latency and energy for neuron core mapping
GPU-Accelerated Hypergraph Partitioning and Placement to Map SNNs on Neuromorphic Hardware
Abstract: SNNs running on neuromorphic hardware use spikes to achieve sparse and energy-efficient communication over a mesh of cores. In turn, system performance heavily depends on the assignment of neurons to cores: the mapping. Since hardware features inter-core multicast and intra-core replication of spikes, we model SNNs as hypergraphs to exploit both opportunities for reducing communication traffic. Mapping thus comprises two NP-hard problems: hypergraph partitioning and placement on the lattice of cores. High-quality solutions to both are critical, yet increasingly difficult as networks scale to millions of neurons. Therefore, we propose a GPU-accelerated pipeline for SNN mapping: a multi-level partitioning scheme is devised around hardware constraints, while placement is initialized through recursive bisection, followed by refinement pulling together strongly connected cores through repeated swaps. Model-based experiments show upwards of 16% lower latency and 42% lower energy for spike movements over existing sequential tools, while our parallel mapper is on average 18-280x faster.