Papers for

processor architects

Papers whose findings have a practical use for this group, as judged from the abstract. Open a paper to read what it means in practice.

Hardware and software improve indirect branch prediction in interpreters

Improving Indirect Branch Prediction in Interpreters via Hardware/Software Co-Design

Abstract: Interpreters have a large indirect-branch footprint, requiring large predictor capacity for accurate prediction. We propose a hardware/software co-design in which a hardware lookahead engine, running ahead of the pipeline with software-provided bytecode metadata, supplies interpreter dispatch targets to the frontend. The engine requires only 1.3 KB of on-chip storage and changes to about 50 lines of CPython code. On 15 CPython server workloads, a 14 KB ITTAGE augmented with the engine reduces bytecode jump MPKI by 73.7% relative to a 16 KB ITTAGE baseline, yielding a 3.2% harmonic-mean IPC speedup.

Mon 28 SeptHardware Architecture
The gist
Interpreters spend a lot of time deciding where to jump next in a program, which can slow down execution. The authors created a system where hardware looks ahead using hints from software about upcoming bytecode instructions, helping the processor guess jump targets more accurately. This approach uses only a small amount of additional hardware memory and small changes to existing Python interpreter code. The result is a large reduction in mispredicted jumps and a modest speedup in processing many server tasks.
Open → 2609.34263v1

Tempo improves processor instruction ordering with lightweight tags

TEMPO: A Tag-Based Framework for Efficient Memory Ordering

Abstract: Weak-memory processors rely on ordering instruc- tions for correctness, yet conventional implementations often en- force them more conservatively than the memory model requires. This over-enforcement manifests as drain-induced retirement stalls at ordering instructions and conservative squash/replay of speculative loads, suppressing legal executions and reducing throughput. We present TEMPO, a tag-based framework for precise microarchitectural implementation of ordering instructions. TEMPO assigns lightweight ordering tags to instructions and decomposes enforcement across retirement-time predicates and completion-time store ordering, allowing the core to enforce required ordering without conservative retirement serialization. TEMPO eliminates unnecessary retirement serialization at ordering instructions and speculative-load squash/replay. In our evaluation, TEMPO reduces geometric-mean normalized exe- cution cycles by 7.9% on native four-thread workloads and improves geometric-mean IPC by 15.9% on an instrumented SPEC2017 dynamic binary translation (DBT) proxy for cross- ISA execution (e.g., x86-on-Arm), while adding only 262 bytes per core.

Sat 19 SeptHardware Architecture
The gist
Processors need to carefully order instructions to work correctly, but current designs often do this more strictly than necessary, causing slowdowns. The authors propose TEMPO, which uses small tags to track instruction order precisely, reducing unnecessary delays and re-executions. This approach helps the processor run tasks faster by enforcing only the required order and improves efficiency without adding much complexity. Tests show it speeds up certain workloads by about 8% in execution time and over 15% in some cross-architecture runs.
Open → 2609.22743v1

Secure speculation design improves software isolation with CHERI

SCHERI: Provably Secure Speculation Under the Constant-Time Policy for CHERI (Extended Version)

Abstract: Capability-based architectures such as CHERI provide strong support for the architectural isolation of software components. To additionally protect against microarchitectural leakage, software can be written in a constant-time fashion. Modern processors, however, rely heavily on speculative execution, which can invalidate the constant-time guarantees and leak isolated secrets transiently. In this work, we show that providing secure speculation for CHERI is non-trivial, and that existing proposals fail to preserve the confidentiality guarantees. We develop a formal framework for reasoning jointly about capability safety, speculative execution, and information-flow security, and use it to demonstrate potential leaks. We then present SCHERI, a new processor design within this framework, and formally prove that it provides end-to-end secure speculation guarantees for the constant-time policy. Our results provide formal foundations and practical guidance for building future capability-based processors, which are resilient to Spectre attacks for constant-time programs.

Tue 15 SeptCryptography and SecurityHardware Architecture
The gist
Speculative execution in modern processors can accidentally reveal secret information, even when software is designed to keep things secret. The authors show that securing speculation on CHERI, a special processor architecture that keeps software parts well-isolated, is difficult and past solutions did not fully protect secrets. They created a new method and processor design called SCHERI that formally guarantees secure speculation while maintaining these secrecy rules. This means software running on SCHERI can better protect secrets from attacks that exploit processor tricks.
Open → 2609.17399v1