Generative Augmentation of Raman Spectra for Glioma Classification

2026-07-11Machine Learning

Machine Learning
AI summary

The authors studied how to improve machine learning for diagnosing glioma tumors using Raman spectroscopy, where only small and varied datasets are available. They created a special AI model that makes fake data similar to real tumor data to help train other models better. While training only on fake data didn’t work as well as training on real data, combining fake and real data led to better results. They also explored a new method to predict tumor types by seeing how well data can be reconstructed. Overall, their work suggests that generating synthetic data can help improve machine learning when real biomedical data is scarce.

Raman spectroscopygliomamachine learningdeep generative modelsvariational autoencoderdata augmentationIDH-status classificationmethylation subtypesynthetic datacross-validation
Authors
Andrei Iuşan, Iulian Vasile, Daria Voiculescu, Ion Petre, Andrei Păun, Bogdan Oancea, Mihaela Păun
Abstract
Access to sufficiently large biomedical datasets remains a major obstacle for machine learning in Raman spectroscopy-based diagnostics. In particular, for glioma analysis, datasets are typically small and heterogeneous, affected by acquisition-specific variability. This work investigates the utility of deep generative augmentation in such a small-cohort setting. We analyze glioma biopsy spectra acquired from 58 tumor samples and consider both binary IDH-status classification and 6-class methylation subtype classification problems. To address the limited size and imbalance of the dataset, we develop a conditional variational autoencoder ($β$-CVAE) capable of generating class-conditioned synthetic Raman spectra. The generated data are evaluated in Train-on-Synthetic, Test-on-Real (TS/TR) and Train-on-Synthetic+Real, Test-on-Real (TSR/TR) settings under a strict patient-isolated cross-validation protocol. Models trained exclusively on synthetic data underperform models trained on real spectra, indicating a substantial domain gap between synthetic and real distributions. However, augmenting the real training data with synthetic spectra consistently improves classification performance across multiple models. These findings indicate that, even with a limited number of independent patient samples, generative models can capture sufficient structure to provide useful regularization for downstream classifiers. We also investigate a reconstruction-based inference strategy, termed Classification by Reconstruction (CbR), in which class prediction is based on reconstruction error under different class conditions. Overall, the results support the use of deep generative augmentation as a practical strategy for improving machine learning robustness in Raman spectroscopy applications characterized by limited biomedical datasets.