Payne-hanek range reduction method made faster and more accurate
A performance enhancement of the Payne-Hanek range reduction algorithm
Mathematical Software
Summary
When computers calculate sine or cosine for very large numbers, they first simplify the number through a process called range reduction. This step can be slow and complicated, especially for big numbers. The authors improved an existing method called Payne-Hanek by removing slow parts like branching and complex integer math, relying only on fast floating-point math instead. Their new approach works efficiently on large inputs and keeps answers very accurate, making trigonometric calculations faster.
What this means in practice
- •For graphics engine developers: Enhance performance of trigonometric computations on large inputs for smoother and faster graphics rendering in games and simulations.
- •For signal processing engineers: Speed up accurate sine and cosine calculations for high-frequency or large-scale signal analysis without sacrificing precision.
Authors
Tue Ly
Abstract
Range reduction plays a crucial role in the accuracy and performance of evaluating trigonometric functions, and is often the primary bottleneck for large floating-point inputs. While fast algorithms such as Cody--Waite work efficiently over narrow intervals, the Payne--Hanek algorithm remains the standard technique for accurate reduction across large floating-point inputs. However, existing implementations of Payne--Hanek suffer from high latency due to heavy branching, conversion overheads, and the use of multi-word integer arithmetic, which hinders SIMD vectorization. In this paper, we analyze and present a branch-free variation of the Payne--Hanek algorithm using only floating-point arithmetic. Our method operates directly over large double-precision inputs ($|x| \ge 2^{16}$) and is well suited to hardware with FMA instructions. We formulate the precision constraints in terms of a truncation error budget, construct a compact lookup table indexed by the input exponent, and prove that the reduced argument is accurate to within one ulp for every input. The same routine can serve both as the complete range reduction of a single-stage implementation and as the fast path of a correctly rounded one, and it improves both latency and throughput over existing implementations. The algorithm is currently implemented in the LLVM libc project.