Hardware and software improve indirect branch prediction in interpreters

Improving Indirect Branch Prediction in Interpreters via Hardware/Software Co-Design

Hardware Architecture

Summary

Interpreters spend a lot of time deciding where to jump next in a program, which can slow down execution. The authors created a system where hardware looks ahead using hints from software about upcoming bytecode instructions, helping the processor guess jump targets more accurately. This approach uses only a small amount of additional hardware memory and small changes to existing Python interpreter code. The result is a large reduction in mispredicted jumps and a modest speedup in processing many server tasks.

What this means in practice

  • For interpreter developers: Add a hardware-software mechanism to reduce mispredicted indirect branches in interpreters for better performance on server workloads.
  • For processor architects: Incorporate hardware lookahead using bytecode metadata to improve branch prediction efficiency and reduce predictor storage requirements.

Authors

Linfeng Zheng, Hiroshi Sasaki

Abstract

Interpreters have a large indirect-branch footprint, requiring large predictor capacity for accurate prediction. We propose a hardware/software co-design in which a hardware lookahead engine, running ahead of the pipeline with software-provided bytecode metadata, supplies interpreter dispatch targets to the frontend. The engine requires only 1.3 KB of on-chip storage and changes to about 50 lines of CPython code. On 15 CPython server workloads, a 14 KB ITTAGE augmented with the engine reduces bytecode jump MPKI by 73.7% relative to a 16 KB ITTAGE baseline, yielding a 3.2% harmonic-mean IPC speedup.