AUTHOR=Chorlton Samuel D. TITLE=Ten common issues with reference sequence databases and how to mitigate them JOURNAL=Frontiers in Bioinformatics VOLUME=4 YEAR=2024 URL=https://www.frontiersin.org/journals/bioinformatics/articles/10.3389/fbinf.2024.1278228 DOI=10.3389/fbinf.2024.1278228 ISSN=2673-7647 ABSTRACT=

Metagenomic sequencing has revolutionized our understanding of microbiology. While metagenomic tools and approaches have been extensively evaluated and benchmarked, far less attention has been given to the reference sequence database used in metagenomic classification. Issues with reference sequence databases are pervasive. Database contamination is the most recognized issue in the literature; however, it remains relatively unmitigated in most analyses. Other common issues with reference sequence databases include taxonomic errors, inappropriate inclusion and exclusion criteria, and sequence content errors. This review covers ten common issues with reference sequence databases and the potential downstream consequences of these issues. Mitigation measures are discussed for each issue, including bioinformatic tools and database curation strategies. Together, these strategies present a path towards more accurate, reproducible and translatable metagenomic sequencing.