Genetic frame selection improves view reconstruction efficiency

Less Is More: Genetic Frame Selection for Efficient Novel View Synthesis

Computer Vision and Pattern Recognition

Summary

When creating new views of a scene from many photos, using all of them isn't always best because some pictures add extra work or lower quality. The authors developed a fast way to pick the most useful photos without fully rebuilding the scene first. They trained a smart system to choose these shots by learning from a slow but thorough search method. Their method picks better photo sets that help create clearer and faster new views, even working well across different scene reconstruction methods and setups.

What this means in practice

  • For computer vision engineers: Improve efficiency by automatically selecting the most informative input views for real-time 3D scene reconstruction without per-scene tuning.
  • For augmented reality developers: Use fewer but better camera frames to generate accurate novel views for AR experiences, reducing computational load on mobile devices.

Authors

Diego E. Farchione, Ramzi Idoughi, Alberto Jaspe-Villanueva, Peter Wonka

Abstract

Feed-forward novel view synthesis reconstructs a scene from many input images in a single forward pass, yet more views do not necessarily improve performance: redundant or poorly chosen frames increase computational cost and may degrade reconstruction quality. We address the problem of selecting, from an already captured sequence, a fixed-size subset of input views that is most informative for reconstructing specified target viewpoints. We propose a render-free view selector that scores candidate frames based on three complementary criteria: target-view coverage, measured against observed frames that stand in for the targets, redundancy with previously selected views, and image sharpness. A lightweight scoring network then selects the most informative frames without rendering, reconstruction, or per-scene optimization at inference time. To train the selector, we distill an expensive offline search procedure in which a genetic algorithm identifies high-quality subsets by directly optimizing reconstruction performance on training scenes. The selector learns to reproduce these choices from geometric and image-level features alone. Across six datasets and multiple input budgets, our method consistently outperforms both geometric and reconstruction-aware view-selection baselines while incurring significantly lower selection costs than reconstruction-based alternatives. Moreover, carefully selected subsets can outperform feed-forward reconstruction from the full input sequence. The learned selector generalizes across diverse reconstruction paradigms (feed-forward, 3D Gaussian Splatting, and NeRF), to object-targeted reconstruction and to a cross-capture setting in which the target views come from a separate acquisition pass. More broadly, our results indicate that explicitly reasoning about target relevance and inter-view redundancy is a fundamental factor in efficient scene reconstruction.