Papers for
processor architects
Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.
Hardware and software improve indirect branch prediction in interpreters
Improving Indirect Branch Prediction in Interpreters via Hardware/Software Co-Design
Abstract: Interpreters have a large indirect-branch footprint, requiring large predictor capacity for accurate prediction. We propose a hardware/software co-design in which a hardware lookahead engine, running ahead of the pipeline with software-provided bytecode metadata, supplies interpreter dispatch targets to the frontend. The engine requires only 1.3 KB of on-chip storage and changes to about 50 lines of CPython code. On 15 CPython server workloads, a 14 KB ITTAGE augmented with the engine reduces bytecode jump MPKI by 73.7% relative to a 16 KB ITTAGE baseline, yielding a 3.2% harmonic-mean IPC speedup.
Tempo improves processor instruction ordering with lightweight tags
TEMPO: A Tag-Based Framework for Efficient Memory Ordering
Abstract: Weak-memory processors rely on ordering instruc- tions for correctness, yet conventional implementations often en- force them more conservatively than the memory model requires. This over-enforcement manifests as drain-induced retirement stalls at ordering instructions and conservative squash/replay of speculative loads, suppressing legal executions and reducing throughput. We present TEMPO, a tag-based framework for precise microarchitectural implementation of ordering instructions. TEMPO assigns lightweight ordering tags to instructions and decomposes enforcement across retirement-time predicates and completion-time store ordering, allowing the core to enforce required ordering without conservative retirement serialization. TEMPO eliminates unnecessary retirement serialization at ordering instructions and speculative-load squash/replay. In our evaluation, TEMPO reduces geometric-mean normalized exe- cution cycles by 7.9% on native four-thread workloads and improves geometric-mean IPC by 15.9% on an instrumented SPEC2017 dynamic binary translation (DBT) proxy for cross- ISA execution (e.g., x86-on-Arm), while adding only 262 bytes per core.
Secure speculation design improves software isolation with CHERI
SCHERI: Provably Secure Speculation Under the Constant-Time Policy for CHERI (Extended Version)
Abstract: Capability-based architectures such as CHERI provide strong support for the architectural isolation of software components. To additionally protect against microarchitectural leakage, software can be written in a constant-time fashion. Modern processors, however, rely heavily on speculative execution, which can invalidate the constant-time guarantees and leak isolated secrets transiently. In this work, we show that providing secure speculation for CHERI is non-trivial, and that existing proposals fail to preserve the confidentiality guarantees. We develop a formal framework for reasoning jointly about capability safety, speculative execution, and information-flow security, and use it to demonstrate potential leaks. We then present SCHERI, a new processor design within this framework, and formally prove that it provides end-to-end secure speculation guarantees for the constant-time policy. Our results provide formal foundations and practical guidance for building future capability-based processors, which are resilient to Spectre attacks for constant-time programs.