Open code enables new antibody editing method with more flexible mutations
When Edit Flows are Edit Jumps: replicating Edit Flows and EvoFlows
Machine Learning
Summary
Improving antibodies often requires making small, precise changes to their sequences, but previous computer models had limits in how these changes could be made. The authors recreated two earlier, unpublished methods and found they work by making one change at a time at varying speeds. They introduced EditJumps, the first openly available tool that can edit antibodies in a more general way without needing retraining for each family. Their work shows that sharing code is crucial for understanding and fairly comparing methods in this field.
What this means in practice
- •For biotech development teams: Generate diverse antibody variants by applying a flexible editing model without retraining for each antibody family, accelerating lead optimization.
- •For bioinformatics engineers: Use an open-source continuous-time edit framework to build or improve sequence editing tools for protein engineering and related applications.
Authors
Gabriel Bénédict, Melanie Buechler, Gerard Riera-Solà, Chloé de Ancos, Yves Gaetan Nana Teukam, Moritz Freidank
Abstract
Antibody lead optimization calls for a small, bounded set of edits to an existing candidate: substitutions, but also insertions and deletions. Edit-based generative models are the only ones that allocate such an edit budget without fixing the edit positions, the edit count, or the output length in advance. However, the existing approaches Edit Flows and EvoFlows did not release code or complete training specifications. Here, we show that both methods follow the same underlying process -- edits firing one at a time, at learned rates, in continuous time -- the pure-jump case of generator matching over finite sequences. With EditJumps we introduce the first open implementation of this framework, with a single generalist antibody editor trained on 1.66M Observed Antibody Space homolog pairs to propose homolog-like variants of a seed sequence, editing unseen leads zero-shot, without the per-family retraining original approaches require. Replicating this system from scratch exposes why open code is essential for generative biology: reconciling published edit distributions required reverse-engineering an undocumented rate-scaling hyperparameter that dictates realized mutation counts. Moreover, we show that published evaluation metrics are highly sensitive to reference sample size, frequently flipping method rankings. We release our full codebase, automated test suite, and configurations at: https://github.com/VisiumCH/editjumps