RIBOSPAN: A Long-Context RNA Foundation Model for Versatile RNA Modeling
2026-08-24 • Machine Learning
Machine Learning
AI summaryⓘ
The authors created RIBOSPAN, a large RNA model that can understand very long RNA sequences—up to 10,240 nucleotides—much longer than previous models. It uses advanced techniques to focus on every single nucleotide, allowing detailed analysis of entire RNA molecules. Tests show that RIBOSPAN performs well at reconstructing sequences, understanding RNA types, and handling long RNA contexts better than earlier models. They also built a method to design complete messenger RNAs, including ways to optimize parts that don't change the protein made from the RNA. Overall, the authors demonstrate a powerful tool for studying and designing long RNAs.
RNA foundation modelmessenger RNA (mRNA)self-attentionsequence tokenizationnucleotide reconstructionlong-context modelingprotein-coding sequence (CDS)discrete diffusionsequence representationsynonymous codons
Authors
Ziyuan Wang, Bohao Tang, Fei Zhang, Shuo Han, Pengfei Liu
Abstract
Full-length RNAs, particularly messenger RNAs, often exceed the context lengths used to pretrain existing RNA foundation models, limiting complete-transcript modeling at single-nucleotide resolution. We present RIBOSPAN, a 1.61-billion-parameter bidirectional RNA foundation model natively pretrained with context lengths up to 10,240 nt. RIBOSPAN combines dense bidirectional self-attention, single-nucleotide tokenization, and attention-isolated sequence packing to enable high-resolution modeling of complete long RNAs. We evaluate the model through nucleotide reconstruction, a controlled long-context representation benchmark, and frozen RNA-type representation analysis. Native 10K pretraining preserves strong reconstruction at 10,240 tokens, while continued pretraining with 40% masking improves recovery under heavy corruption while preserving representation quality. The long-context benchmark further shows that native 10K models maintain strong contextual responsiveness and context-specific representation separation while keeping perturbation-induced representation changes highly localized. Inference-time YaRN scaling recovers much of the contextual organization lost by direct extrapolation of short-context models, but induces substantially greater distal representation diffusion. Frozen-representation evaluations further demonstrate state-of-the-art RNA representation quality, with RIBOSPAN achieving the strongest overall performance across diverse RNA types and retaining a clear advantage on long RNAs. Building on the same backbone, we develop a multidimensionally conditioned discrete-diffusion framework for full-length mRNA generation and redesign, including synonymous-codon diffusion for protein-preserving CDS optimization. Together, RIBOSPAN establishes a powerful long-context foundation for transferable RNA representation learning and full-transcript mRNA design.