The Unified Evaluation App for DNA Data Storage Codecs
2026-08-10 • Emerging Technologies
Emerging TechnologiesSoftware Engineering
AI summaryⓘ
The authors created an open-source tool to fairly compare different methods of storing data in DNA. This tool tests how well each method works by looking at speed, accuracy, and cost using a shared set of rules. They found that no single method is best in every way, with trade-offs between storage efficiency, error handling, and resource use. Their work helps researchers choose the right approach and move DNA data storage closer to real-world use.
DNA data storageencodingdecodingerror correctionbenchmarkinginformation densitycomputational efficiencydata archivalcodecthroughput
Authors
Aleksandar Anžel, Chisom Anyabolu, Leon Wimbes, Luca Staus, Ihsan Tri Heldian, David Sonnabend, Khawla Elhadri, Samuel Becker, Felix Klein, Raffael Schoen, Michael Schwarz, Marius Welzel, Bernd Freisleben, Dominik Heider, Georges Hattab
Abstract
Background: Deoxyribonucleic acid (DNA) data storage is a paradigm with great potential for ultra-dense and durable information preservation. However, the rapid proliferation of coding schemes, or codecs, each with their own design constraints and reporting practices, has led to a fragmented landscape that lacks a standardized comparative assessment. Methods: We developed an open-source, modular benchmarking platform that systematically integrates and evaluates state-of-the-art DNA storage encoding and decoding methods (codecs). Our approach uses a curated, diverse set of baseline data and applies multidimensional assessment criteria that are aligned with the consensus standard of the DNA Data Storage Alliance. These criteria include encoding/decoding throughput, computational efficiency, error correction performance across substitutions, insertions, and deletions, and cost efficiency. Results: The developed platform integrates standardized wrapper functions for encoding and decoding, allows for the integration of new methods, and automates reproducible evaluations with comprehensive visual and tabular reporting. Benchmarking both contemporary and classical codecs using their default parameters and multiple metrics demonstrates that no single algorithm is optimal across all evaluated dimensions. The trade-offs between information density, success rate, runtime, and cost are quantified and shown to be critical factors in the design of future-proof formats. Conclusions: Our work establishes a rigorously standardized, open-source evaluation framework that enables reproducible benchmarking, supports evidence-based codec selection, and provides the necessary foundation for translating DNA data storage from experimental research into deployable archival systems.