Papers for

high-performance computing developers

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Python framework enables fast quantum computing across CPUs GPUs and FPGAs

Python in the front, party in the Backline: compiling quantum workloads across CPUs, GPUs, and FPGAs

Abstract: Moving from quantum research and development to production-grade, fault-tolerant quantum workload execution remains one of the most significant challenges facing quantum platform builders. While Python frameworks have enabled an easy entry point for quantum algorithm design, the low-latency requirements for real-time quantum error correction (QEC) demand performance that traditional interpreted environments cannot provide. FPGAs and ASICs play a central role at these layers, but their specialized programming models make development rigid and time-consuming. CPUs, GPUs, and other accelerators introduce a different challenge: as infrastructure becomes increasingly heterogeneous, programming across different devices and their associated abstractions becomes more complex. Allowing researchers to write workloads in high-level languages that map to low-latency execution across diverse distributed target platforms will enable the development of key infrastructure for utility-scale quantum systems. For this, we introduce $\textit{Backline}$, a heterogeneous compilation and runtime framework built within PennyLane and Catalyst. Backline allows us to design and build quantum-classical workloads for high-performance and low-latency devices, with compilation directly from a Python interface through MLIR. We demonstrate the compilation and execution of several quantum workloads with low-latency data movement across a mix of CPUs, GPUs, and FPGAs, for both local and distributed remote hardware targets, all from a vendor-agnostic Python frontend. With an AMD VPK120 FPGA board as the controller, issuing each round from its hardware-handshake engine, we measured median steady-state round-trip latencies over RoCE v2 of $2.305~μ$s to an AMD Ryzen Threadripper PRO CPU and $4.5~μ$s to an AMD Instinct MI210 GPU across $10^6-1$ rounds per path, demonstrating microsecond-scale synchronous co-processing.

Tue 8 SeptDistributed, Parallel, and Cluster ComputingProgramming Languages
The gist
Quantum computers need both smart software and fast hardware to run complex calculations reliably. The authors created Backline, a tool that lets programmers write quantum tasks in Python, which then run quickly on different hardware like CPUs, GPUs, and FPGAs. This helps combine easy coding with the speed needed for real-time quantum error correction. They showed their system could send and receive data in just a few microseconds between devices, which is important for future quantum machines.
Open 2609.09270v1

New sample-based method speeds exact Top-K sparse attention on GPU

Sample-Guided Exact Top-K Selection for Long-Context Sparse Attention

Abstract: Sparse attention bounds downstream attention work by retaining a fixed-size subset of indexed tokens, but its standalone exact Top-$K$ stage must still process materialized score rows whose length grows with context. Production radix selectors discover their first actionable boundary only after a complete-row pass, forcing another row-scale traversal before exact refinement. We observe that locating a compact upper tail requires substantially less resolution than identifying the exact rank boundary, and that fixed-stride partial views of the current row remain calibrated to the corresponding complete-row rank across ragged lengths. We present HPC-Ops Top-K, a sample-guided exact selector for ragged sparse-attention score rows. A fixed-stride view proposes a row-local coarse boundary; the mandatory complete-row pass certifies its sufficiency, forms the admitted candidate set, and initializes exact FP32 refinement over the unresolved frontier. A nested secondary boundary and exact recovery handle underfilled proposals before any output is committed, so sampling controls common-path work but never correctness. The GPU implementation fuses complete-row certification and candidate formation, and combines persistent, KV-split, and direct-exact execution behind graph-capturable ragged-row dispatch. We evaluate HPC-Ops Top-K on indexer scores from Hy4-Preview. It outperforms the fastest verified external exact baseline by $1.29$--$1.75\times$ across 20 operator configurations, with a $1.55\times$ geometric-mean speedup. It further achieves $1.36\times$ and $1.48\times$ speedups on two framework-derived sparse-attention traces. The implementation is available in HPC-Ops, Tencent's open-source high-performance operator library for LLM inference, at https://github.com/Tencent/hpc-ops.

Tue 8 SeptDistributed, Parallel, and Cluster Computing
The gist
Large language models use complex attention mechanisms that need to select the most important tokens from a long context quickly. The authors found that looking at a small, spaced sample of data first can guide the exact selection process more efficiently without losing accuracy. They implemented this method on GPUs and showed it is about 1.3 to 1.7 times faster than previous exact methods across many test cases. This improvement helps make large models work faster during inference.
Open 2609.08450v1