Algorithm optimizes verification of AI answers within limited resources

Test-Time Scaling via Budgeted Multi-Attribute Verification

Artificial Intelligence

Summary

Checking AI answers carefully can take a lot of computing power, especially when you want to confirm different qualities of each answer. This paper shows how to decide which answers to check and what qualities to focus on while keeping within a set budget. The authors created an algorithm named BMA-GAI that smartly spends effort to confirm as many good answers as possible without extra waiting steps. Their method is proven to be very efficient and better at finding good answers than other ways, based on tests with fake data and real AI answer checking.

What this means in practice

  • For ai system engineers: Manage computational budget when verifying large language model answers by prioritizing which answers and attributes to check.
  • For quality assurance teams: Certify the quality of AI-generated content more efficiently by adapting verification efforts along multiple quality aspects.

Authors

Bo Xue, Ji Cheng, Shen-Huan Lyu, Yuanyu Wan, Shuang Qiu

Abstract

Verifying LLM-generated answers under a shared computational budget requires jointly deciding which candidates to inspect and which verification attributes to evaluate. We formulate this problem as multi-attribute good-arm identification under a global budget: each candidate is an arm evaluated along several costly attributes, and the goal is to certify as many candidates as possible whose mean scores exceed the prescribed thresholds on all attributes. We propose \textsc{BMA-GAI}, an algorithm that combines cost-aware arm selection with adaptive sampling of attributes. Every observation serves both to guide adaptive allocation and to support anytime-valid certification, which removes the need for a separate confirmation stage. We establish an asymptotic coverage guarantee for \textsc{BMA-GAI} and derive a matching information-theoretic converse that characterizes the intrinsic complexity of the problem, thereby proving that \textsc{BMA-GAI} is first-order optimal away from critical budget levels. Experiments on synthetic benchmarks and an LLM answer-verification task show that \textsc{BMA-GAI} allocates the verification budget more efficiently and certifies more high-quality candidates than competing methods.