Plos iconPlosSep 3, 2026 ~1 min source read

Balanced DNA interpolation improves learning of genetic distance-informed embeddings in plants

These augmentations do not take into account the inherent structures of populations and species, potentially blurring the lines between entities within genetic datasets. Here, we propose DNAInterpolator, an approach based on interpolation of DNA sequences within a given dataset that presents a neighbor-guided alternative to random mutations.

Balanced DNA interpolation improves learning of genetic distance-informed embeddings in plants

Share this story

Send the public story page.

Useful takeaways from this story.

Our new approach leverages the potential of taxonomic DNA datasets for modern machine learning applications.

Genetic data augmentation usually refers to the introduction of random base variations, translocations, and reverse complementing.

These augmentations do not take into account the inherent structures of populations and species, potentially blurring the lines between entities within genetic datasets.

Building the complete brief

The page is ready to read now. The fuller skim-friendly version will appear here automatically.

The useful part

Genetic data augmentation usually refers to the introduction of random base variations, translocations, and reverse complementing. These augmentations do not take into account the inherent structures of populations and species, potentially blurring the lines between entities within genetic datasets. Here, we propose DNAInterpolator, an approach based on interpolation of DNA sequences within a given dataset that presents a neighbor-guided alternative to random mutations.

How it works

  • We tested interpolation as an augmentation technique using four flowering plant datasets and an artificial neural network trained to predict genetic distances between paired samples.
  • To address unequally distributed distances within our training datasets, we examined the effect of balancing the distance distribution by curating interpolated sequences.
  • Our new approach leverages the potential of taxonomic DNA datasets for modern machine learning applications.
  • Although the amount of DNA data needed to train state-of-the-art machine learning models often exceeds what can realistically be collected and sequenced in biological studies, the number of samples can be...
  • We found that balancing helps models capture genetic distances across the full distance range by strengthening performance in underrepresented regions of the distribution.

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app