Combining retrieval methods improves Taiwanese history question answers

Bridging Static and Agentic RAG for Taiwanese Historical Question Answering

Computation and LanguageInformation Retrieval

Summary

Answering history questions about Taiwan can be done using different methods to find helpful information. One method changes its search based on what it finds, while another uses a fixed plan. The authors found neither method is best for all questions; instead, they work best when combined and their strengths are chosen question by question. They built a system that picks the better answer each time, improving overall results. This shows that mixing search strategies and choosing between them can help answer questions more accurately.

What this means in practice

  • For historical data teams: Improve accuracy in answering Taiwanese historical questions by combining adaptive and static search methods to select the best response per query.
  • For digital library developers: Integrate dual retrieval strategies with a decision system to enhance user experience when querying cultural or historical archives.

Tested on one dataset.

Authors

Kai-Hsin Chen, Wei-Yu Chen, Xuanjun Chen, Jyh-Shing Roger Jang

Abstract

Agentic retrieval-augmented generation (RAG) enables language models to adapt retrieval based on previously retrieved evidence, but it remains unclear whether such adaptive orchestration consistently outperforms well-designed static pipelines. We conduct a controlled comparison of agentic and static RAG for Taiwanese historical question answering, sharing the same generator and hybrid retrieval backend. Despite similar aggregate performance, the two pipelines differ on 70.83% of questions, with their advantages largely canceling out when averaged. An oracle that selects the better response per question improves the composite score by 0.2417 over the better individual pipeline, revealing substantial headroom for question-level selection. We therefore introduce a post-hoc selector that compares the two responses and their cited evidence, significantly outperforming either individual pipeline and recovering 60.34% of the oracle headroom. These results show that aggregate comparisons can obscure meaningful question-level differences between retrieval strategies, suggesting that exploiting their complementarity may be more fruitful than seeking a universally superior pipeline.