Tempo improves processor instruction ordering with lightweight tags

TEMPO: A Tag-Based Framework for Efficient Memory Ordering

Hardware Architecture

Summary

Processors need to carefully order instructions to work correctly, but current designs often do this more strictly than necessary, causing slowdowns. The authors propose TEMPO, which uses small tags to track instruction order precisely, reducing unnecessary delays and re-executions. This approach helps the processor run tasks faster by enforcing only the required order and improves efficiency without adding much complexity. Tests show it speeds up certain workloads by about 8% in execution time and over 15% in some cross-architecture runs.

What this means in practice

  • For processor architects: Design processor cores that enforce memory ordering more precisely to reduce execution stalls and improve throughput.
  • For compiler and runtime developers: Optimize cross-ISA dynamic binary translation systems by reducing costly instruction ordering overheads on target architectures.

Authors

Pranith Kumar, Prasun Gera, Hyojong Kim, Chulhyung Park, Hyesoon Kim

Abstract

Weak-memory processors rely on ordering instruc- tions for correctness, yet conventional implementations often en- force them more conservatively than the memory model requires. This over-enforcement manifests as drain-induced retirement stalls at ordering instructions and conservative squash/replay of speculative loads, suppressing legal executions and reducing throughput. We present TEMPO, a tag-based framework for precise microarchitectural implementation of ordering instructions. TEMPO assigns lightweight ordering tags to instructions and decomposes enforcement across retirement-time predicates and completion-time store ordering, allowing the core to enforce required ordering without conservative retirement serialization. TEMPO eliminates unnecessary retirement serialization at ordering instructions and speculative-load squash/replay. In our evaluation, TEMPO reduces geometric-mean normalized exe- cution cycles by 7.9% on native four-thread workloads and improves geometric-mean IPC by 15.9% on an instrumented SPEC2017 dynamic binary translation (DBT) proxy for cross- ISA execution (e.g., x86-on-Arm), while adding only 262 bytes per core.