AstroSpecLM uses language AI to explain and analyze star light spectra

AstroSpecLM: A Spectrum-Language Model for Evidence-Grounded Astronomical Spectral Analysis

Artificial Intelligence

Summary

Astronomical spectra are patterns of light from space that hold important information, but understanding them usually requires experts. The authors created AstroSpecLM, a system that links these light patterns to a language AI to answer questions and explain its findings in simple language. They first convert complex spectra into key facts, then train the AI to use those facts to have conversations about the data. Their system matches specialized methods for identifying celestial objects and measuring distances while also giving clear, detailed explanations. This shows it’s possible to build AI models that both predict and explain scientific data like star light spectra.

astronomical spectralanguage modelDESI spectraspectral classificationredshift estimationinstruction-followingmachine learningspectral featuresevidence groundingnatural language explanation

Authors

Jinghang Shi, Yanxia Zhang, Ali Luo, Changhua Li, Xiao Kong

Abstract

Astronomical spectra encode rich physical information, but drawing scientific conclusions from spectral features typically requires expert interpretation. This paper presents AstroSpecLM, a spectrum-language model that connects one-dimensional DESI spectra with Qwen3-4B to answer questions and provide explanations grounded in spectral evidence. Instead of generating question-answer pairs directly from templates or raw catalog fields, we first distill each spectrum into a compact set of catalog- and spectrum-derived facts, then use these facts as references to generate instruction-following conversations. The resulting model is competitive with specialist supervised baselines on classification and redshift estimation, while additionally producing natural-language explanations that reference specific spectral features. Our results indicate that grounding a language model in one-dimensional scientific spectra is feasible, and that fact-mediated instruction data yields a model capable of both prediction and explanation.