Context projection matches full results while cutting memory use

Completed Pairs Hide Capped Failures: A ReVerPi Case Study of Selective Context Projection

Artificial Intelligence

Summary

The paper studies a technique called context projection that shrinks old tool outputs into small, easy-to-retrieve pieces. The authors compare this method to using full original data in a coding task environment. They find that both ways succeed equally often, but the projected method uses less overall memory while sometimes needing more interaction. Stopping rules and replaying both methods independently matter for fair evaluation. This work focuses on a specific recorded experiment rather than proving general superiority.

What this means in practice

  • For software tool developers: Design debugging assistants that use compact context excerpts to save memory without losing accuracy in code understanding tasks.
  • For ai system evaluators: Set up fair benchmarking pipelines that run both full and projected inputs independently, measuring completion and token use to assess efficiency.

Authors

Guangzhe Zhang

Abstract

Context projection replaces older tool observations with compact, addressable excerpts, reducing repeated input while potentially adding evidence-retrieval turns. We study this trade-off in ReVerPi, a Pi extension with archived observations and matched full/projected continuations. In an 86-run source-reading campaign with 641 model requests, the 15 completed pairs show identical success: 12/15 per arm. Twelve further boundary runs stop, with the runner suppressing the companion whenever the first arm fails to complete. Restoring all 27 boundary runs bounds projected-minus-full success between $-$9 and +1 tasks. One omitted, selector-chosen projected continuation successfully retrieves archive text yet exhausts twelve requests; its full counterpart answers in three. The eleven jointly correct pairs form a fully observed success stratum within this recorded frame: projection reduces aggregate logical tokens by 25%, while increasing the median pair's tokens by 29% and total suffix requests from 35 to 55. Separating fitting from evaluation changes the selector's apparent tie: outside its four fitting pairs, it incurs one extra failure and 8.6% more logical tokens over thirteen comparable runs. This methodological case study connects stopping rules, known bounded failures, unexecuted companions, and resource aggregation. Its findings concern the recorded campaign, rather than population noninferiority or superiority over unrestricted Pi. Evaluations should retain every intervention boundary, execute both allocated arms independently of the first arm's completion, and report completion alongside interaction and token expenditure.