Phylogeny guided method generates ancestral bird sounds with realistic quality
Generating the Unheard: Phylogeny-Guided Latent Generation for Ancestral Sound Reconstruction
Sound
Summary
The researchers developed a way to recreate what ancient birds might have sounded like, using recordings of modern birds. They turned these sounds into a special format, then used the family tree of birds to guess what the sounds of their ancestors would be. Their method creates new, believable bird vocalizations for each ancestor, rather than just copying existing sounds. Tests showed this approach worked well on two different bird families, producing realistic and scientifically consistent ancestral calls.
What this means in practice
- •For wildlife sound archivists: Reconstruct missing ancestral bird calls to enrich audio collections and improve archival completeness.
- •For audio software developers: Integrate phylogeny guided sound generation algorithms to create novel ancestral audio samples for educational and entertainment tools.$Commercial implications: Enables development of new sound design products using scientifically guided ancestral bird call synthesis.
Authors
Tianyi Xu, Shrinaath Narasimhan, Evan Gorstein, Santiago Perea, Yunyi Shen, Claudia Solís-Lemus
Abstract
What did an ancestral bird species sound like? Existing ancestral state reconstruction methods can infer low-dimensional traits such as morphological characters at internal nodes of a phylogenetic tree, but no one has tried to produce rich perceptual signals such as audio. Some of the challenges include inferred representations that are either too low-dimensional to decode or lie in non-generative feature spaces, so no method to date can produce ancestral audio. We introduce the first framework that generates plausible ancestral vocalizations. Our pipeline encodes bird recordings into a VAE latent space, learns a low-dimensional trait projection aligned with phylogenetic distances, performs ancestral inference in this trait space, and recovers decodable latents through an anchored inverse lift before emitting novel waveforms for each ancestral node. Because the entire pipeline stays within a decodable latent space, every internal node receives a genuinely new audio output representing plausible intermediate ancestral sounds unavailable to retrieval-based alternatives. Experiments on two phylogenetically distant bird clades, 21-species Tyrannidae and 19-species Paridae, show that our method is the only approach that simultaneously achieves genuine generation, phylogenetic consistency, and naturalistic audio quality across both datasets.