USA300 has remained the dominant community and healthcare associated methicillin-resistant
Here, we take advantage of the accrued, publicly available data, as well as two newly sequenced pre-epidemic historical isolates from 1996, and a very early diverging ACME-negative NAE genome, to understand the pre-epidemic evolution of USA300. We use database mining techniques to emphasize genomes similar to pre-epidemic isolates, with the goal of reconstructing the early molecular evolution of the USA300 lineage.
Phylogenetic analysis with these genomes confirms that the NAE and SAE USA300 lineages diverged from a most recent common ancestor around 1970 with high confidence, and it also pinpoints the independent acquisition events of the of the ACME and COMER loci with greater precision than in previous studies. We provide evidence for a North American origin of the USA300 lineage and identify multiple introductions of USA300 into South and North America. Notably, we describe a third major USA300 clade (the pre-epidemic branching clade; PEB1) consisting of both MSSA and MRSA isolates circulating around the world that diverged from the USA300 lineage prior to the establishment of the South and North American epidemics. We present a detailed analysis of specific sequence characteristics of each of the major clades, and present diagnostic positions that can be used to classify new genomes.