Hardware and software improve indirect branch prediction in interpreters
Improving Indirect Branch Prediction in Interpreters via Hardware/Software Co-Design
Hardware Architecture
Summary
Interpreters spend a lot of time deciding where to jump next in a program, which can slow down execution. The authors created a system where hardware looks ahead using hints from software about upcoming bytecode instructions, helping the processor guess jump targets more accurately. This approach uses only a small amount of additional hardware memory and small changes to existing Python interpreter code. The result is a large reduction in mispredicted jumps and a modest speedup in processing many server tasks.
What this means in practice
- •For interpreter developers: Add a hardware-software mechanism to reduce mispredicted indirect branches in interpreters for better performance on server workloads.
- •For processor architects: Incorporate hardware lookahead using bytecode metadata to improve branch prediction efficiency and reduce predictor storage requirements.
Authors
Linfeng Zheng, Hiroshi Sasaki
Abstract
Interpreters have a large indirect-branch footprint, requiring large predictor capacity for accurate prediction. We propose a hardware/software co-design in which a hardware lookahead engine, running ahead of the pipeline with software-provided bytecode metadata, supplies interpreter dispatch targets to the frontend. The engine requires only 1.3 KB of on-chip storage and changes to about 50 lines of CPython code. On 15 CPython server workloads, a 14 KB ITTAGE augmented with the engine reduces bytecode jump MPKI by 73.7% relative to a 16 KB ITTAGE baseline, yielding a 3.2% harmonic-mean IPC speedup.