Small reasoning models improve by asking bigger models for help
Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge
Artificial IntelligenceComputation and Language
Summary
Thinking longer doesn’t always help small language models solve hard problems if they lack the right knowledge. The authors found that these models tend to circle around known answers instead of discovering new ones on their own. They created FlyBy, a method where small models try to reason first, then ask bigger, smarter models for help only when needed. This approach improves problem solving while keeping computing costs low.
What this means in practice
- •For ai service providers: Cut costs by using small reasoning models that selectively query larger models only when extra knowledge is needed during problem solving.
- •For software developers in ai: Build efficient multi-model reasoning systems where lightweight models handle most tasks and escalate complex parts to more powerful models dynamically.
Authors
Chanuk Lee, Minki Kang, Sangwoo Park, Woongyeong Yeo, Jinheon Baek, Sung Ju Hwang
Abstract
Scaling test-time computation is a powerful way to improve language-model reasoning, and is particularly appealing for small reasoning models (sRMs) that are cheap to serve. However, is additional thinking always the right operation? By intervening at intermediate reasoning states across two model families and multiple scales, we find that self-refinement largely consolidates probability mass onto solutions already reachable from the current state, rather than making new ones reachable. These interventions reveal two failure regimes: execution bottlenecks, where the correct path is reachable and reflection can recover it, and knowledge bottlenecks, where relevant external information makes it reachable. Motivated by this distinction, we introduce FlyBy, a selective querying framework, and train 4B and 8B variants to reason first, diagnose what remains unresolved, and, at a knowledge bottleneck, query stronger models whose parametric knowledge extends beyond its own. Supervised fine-tuning bootstraps a multi-depth query action, and cost-aware reinforcement learning calibrates whether to query, what to ask, and how much to spend. On 1,158 hard problems across six benchmarks, FlyBy-4B achieves 45.96% pass@8, surpassing Qwen3-14B (41.64%) at 2.7 times lower serving cost, while also exceeding Qwen3-8B in pass@1 (16.85% vs. 15.31%). Scaling to FlyBy-8B further improves pass@8 to 51.81%.