Marker-Delimited Codes for Short-Blocklength, High-Rate Coding over Multi-Read Edit Channels

2026-08-31Information Theory

Information Theory
AI summary

The authors address the problem of reading data stored in DNA, which often comes with errors like wrong, missing, or extra letters. They create a two-part error correction method using a special inner code called marker-delimited code (MDC) and an outer LDPC code to fix mistakes more reliably. Their method helps recover the original data better, especially when handling short pieces of data and high data rates. Overall, the authors improve how DNA data storage deals with common read errors.

DNA data storageedit errorssubstitutionsdeletionsinsertionsmarker-delimited codeLDPC codebelief propagationsoft informationerror correction
Authors
Sinan Ates Yercan, Marc Antonini, Serge Kas Hanna
Abstract
The read process of DNA-based data storage systems generates multiple noisy copies of the stored DNA sequences, affected by edit errors consisting of substitutions, deletions, and insertions. Motivated by the challenge of ensuring reliable data retrieval in the presence of edit errors, we present a concatenated coding scheme that accounts for practical design constraints in DNA storage. We introduce and apply the marker-delimited code (MDC) as the inner code, which enables fast and reliable computation of symbolwise a posteriori probabilities (APPs). We combine MDC with an outer LDPC code. The LDPC is decoded via belief propagation using the soft information generated by MDC. Our results show that, in comparison with prior work, this construction provides more efficient error correction over multi-read edit channels in the short-blocklength and high-rate regime.