Sequence-informed biodiversity metrics for eDNA
From conventional feature tables to taxonomy-free, sequence-informed representations of environmental DNA.Environmental DNA (eDNA) provides a scalable and non-invasive way to survey biodiversity from water, soil, air and other environmental samples. However, conventional metabarcoding workflows usually reduce sequences to amplicon sequence variants (ASVs), operational taxonomic units (OTUs) or taxonomic assignments. These features depend on reference databases and study-specific bioinformatic choices, which can limit reproducibility and make independent datasets difficult to integrate.
This project asks how we can use the molecular information contained in eDNA sequences more directly. I am exploring three complementary representations: pairwise sequence-distance matrices, DNA image representations based on nucleotide or k-mer composition, and vector embeddings generated by DNA language models. Together, these approaches retain relationships among sequences without requiring every sequence to be assigned to a named taxon or a study-specific operational unit.
In a proof-of-concept reanalysis of the Earth Microbiome Project, sequence-informed metrics captured dimensions of diversity that complemented conventional ASV-based measures. Several representations were less dominated by study identity and retained stronger environmental signals, while k-mer and language-model representations performed comparably to ASV feature tables in ecological classification.
By making eDNA datasets more comparable and transferable, sequence-informed approaches could support biodiversity assessment across regions and studies, distinguish ecological signals from methodological variation, and improve prediction for new samples and environments. The broader goal is to make fuller use of the molecular variation already contained in eDNA data.