High-Throughput Multiplex Detection of Antibiotic-Resistant Genes and Virulence Factors in Using Digital Multiplex Ligation Assay
High-Throughput Multiplex Detection of Antibiotic-Resistant Genes and Virulence Factors in Using Digital Multiplex Ligation Assay
6, June 2025
[Link]
From Eawag,* Swiss Federal Institute of Aquatic Science and Technology, Duebendorf, Switzerland; the Department of Biosystems Science and Engineering,y
ETH Zurich, Basel, Switzerland; the Swiss Institute of Bioinformatics,z Lausanne, Switzerland; the Department of Environmental Systems Science,x ETH
Zurich, Zurich, Switzerland; the Department of Biology,{ University of Turku, Turku, Finland; the Swiss Tropical and Public Health Institute,k Allschwil,
Switzerland; and the University of Basel,** Basel, Switzerland
The vast majority of Escherichia coli are harmless com- proteins that enable E. coli to cause disease by affecting a
mensals living within the human gut, but a critical fraction wide range of cellular processes.1
are pathogenic.1 Pathogenic E. coli are among the leading The emergence of antimicrobial resistance (AMR) in
cause of moderate-to-severe diarrhea in low-resource set- pathogenic E. coli strains presents an additional challenge,
tings,2 and diarrheal diseases are the major cause of child- complicating treatment options and leading to increased
hood mortality, with yearly estimated 444,000 deaths in morbidity, mortality, and health care expenditures.4 Low-
children aged <5 years ([Link] and middle-income countries are the most burdened with
health/diarrhoeal-disease, last accessed April 14, 2024). both diarrheal diseases and antimicrobial resistance carriage,
Intestinal pathogenic E. coli strains are classified into six but they often have insufficient capacity to establish and/or
pathotypes: enteroaggregative, enterohemorrhagic,
enteroinvasive, enteropathogenic, enterotoxigenic, and
diffusely adherent E. coli.3 Pathotypes are distinguished Supported in whole or in part by the Swiss National Science Foundation
by virulence factors (VFs), which are small molecules and grant 192763 (T.R.J.).
Copyright ª 2025 Association for Molecular Pathology and American Society for Investigative Pathology. Published by Elsevier Inc.
This is an open access article under the CC BY license ([Link]
[Link]
Conforti et al
maintain surveillance systems,5,6 leaving much of the AMR comprehensive surveillance of potentially hundreds of
burden in these settings unquantified.7 genes across hundreds of bacterial isolates in a single
Surveillance of AMR and pathogenicity is conducted sequencing effort.
through global and regional networks, including clinical and In this study, the application of dMLA was extended
environmental assessments.8 The World Health Organiza- beyond b-lactamase genes to include 43 genes related to
tion Global Antimicrobial Resistance and Use Surveillance antibiotic resistance, VFs, and E. coliespecific genetic
System is a key example, promoting standardized global markers. This expansion aims at the comprehensive char-
collection and sharing of AMR data.9 Beyond clinical acterization of E. coli isolates through identification of
sources, the Global Antimicrobial Resistance and Use Sur- pathotypes, resistance across six antibiotic classes, phy-
veillance System has implemented a One Health approach logroup classification, and species-level confirmation.
with the Tricycle Surveillance Project, which integrates data
from humans, animals, and environmental sources.10
Additional networks include the CDC Gonococcal Isolate
Materials and Methods
Surveillance Program in the United States ([Link] Selection of Target Genes and Retrieval of Reference
[Link]/view/cdc/40081, last accessed March 13, 2025), Sequences
the European Antimicrobial Resistance Surveillance
Network in Europe,11 and the Central Asian and European Antibiotic-Resistant Genes
Surveillance of Antimicrobial Resistance network,12 all Nineteen19 antibiotic-resistant genes were selected on the basis
aligning with the World Health Organization initiatives. The of their prevalence in human-associated environments, their
academic and private sectors also significantly enhance potential for horizontal transfer among bacteria, their presence
AMR surveillance.13 in ESKAPE (Enterococcus faecium, Staphylococcus
AMR surveillance incorporates both phenotypic and aureus, Klebsiella pneumoniae, Acinetobacter
genotypic approaches. Predominantly, the above-mentioned baumannii, Pseudomonas aeruginosa, and Enterobacter spp.)
networks generate culture-based AMR data through pathogens,23 or their clinical threat due to rapid emergence and
phenotypic assays, which evaluate antimicrobial effective- diffusion, as highlighted in multiple studies24e27 (Table 1).
ness against bacteria, assisting in therapeutic guidance and Reference sequences were obtained from the Comprehensive
resistance mapping.14 Nevertheless, culture-based methods Antibiotic Resistance Database ([Link] last
are limited in the range of antibiotics they can test, are time- accessed September 29, 2022), focusing only on homologs
consuming and resource intensive, and cannot distinguish related to E. coli, the target organism. The selection process
between pathogenic and nonpathogenic E. coli strains.15 for downloading specific homologs from the Comprehensive
Therefore, comprehensive E. coli characterization requires Antibiotic Resistance Database involved verifying the
methods beyond culture-based protocols. Genotypic tech- presence of E. coli in either resistomes with perfect matches
niques, such as quantitative real-time PCR, digital PCR, or resistomes with sequence variants. If E. coli was not
multiplex PCR, and sequencing, offer rapid detection of identified in either category, only those homologs were
resistance genes and insights into resistance mechanisms included that demonstrated a prevalence of >1.5% in whole-
and VFs not identifiable by phenotypic tests. However, genome shotgun assemblies within the National Center for
despite their precision, genotypic methods are constrained Biotechnology Information (NCBI) database ([Link]
by their capacity to identify a limited number of targeted [Link], last accessed February 17, 2021). The
DNA regions simultaneously, presenting a significant comprehensive list of chosen homologs is available in
challenge for large-scale surveillance.16e18 Metagenomics Supplemental Table S1.
offers a comprehensive overview of microbial communities,
including nonculturable bacteria, but its high costs, limited Virulence Factors
sensitivity, and data complexity limit its practicality for Sixteen16 VFs associated with five E. coli pathoty-
routine monitoring.19,20 pesdenteroaggregative E. coli, enterohemorrhagic E. coli,
The digital multiplex ligation assay (dMLA) is a low- enteroinvasive E. coli, enteropathogenic E. coli, and en-
cost, high-throughput molecular technique initially terotoxigenic E. colidwere selected (Table 1). Reference
developed by Tamminen et al,21 in 2020, to detect b- sequences for these VFs were retrieved from the Virulence
lactamase genes. This assay offers a cost-effective and Factors Database ([Link] last accessed
efficient approach for screening large isolate collections, January 23, 2021).28 The full data set of verified and pre-
such as those stored in archives, or obtained within the dicted VFs was downloaded from the Virulence Factors
scope of monitoring (eg, nationally representative Database. After filtering of the data set to include only E.
household water quality testing, such as the Multiple In- coli sequences, the VFs of interest were identified by the
dicator Cluster Survey).22 By using in silico probe design, sequence titles, and the corresponding available references
ligation-based detection, and next-generation sequencing, for each VF were downloaded. Between 1 and 15 reference
the multiplexing ability of dMLA significantly reduces sequences were obtained for each VF, with accession
costs and enhances throughput. This enables numbers detailed in Supplemental Table S2. These
Table 1 Target Genes for Antibiotic Resistance, Virulence, and Phylogroup Identification in Escherichia coli Using Digital Multiplex
Ligation Assay
Resistance Antibiotic class Antibiotic-resistant gene
Aminoglycosides aadA, aph(3’)-I, aac(6’)-I, aac(3)-II, aac(3)-VI, ant(200 )-Ia, aph(6)-I
Colistin mcr
Quinolones qnrB, qnrS
Macrolides ermB, mphA
Sulfonamides/diaminopyrimidines dfrA, sul1, sul2, sul3
Tetracycline tetA, tetB, tetM
Pathogenicity Pathotype Virulence factor
EAEC astA, aggR, aatA, aafA, aap, aaiC
EPEC bfpA, bfpF
EIEC invE, ipaH
EPEC/EHEC eae
EHEC stx1, stx2
ETEC eltA, eltB, estIa
Phylogroup Genetic marker
chuA, yjaA, trpA, arpA, TspE4.C2
E. coli positive control uidA
E. coli negative control gapdh
Resistance marker intI1
This table categorizes the selected genes into three main sections: antibiotic resistance genes across six antibiotic classes, virulence factors for five E. coli
pathotypes, and genetic markers for E. coli phylogroup identification, including positive and negative controls for E. coli detection and a resistance marker.
EAEC, enteroaggregative E. coli; EHEC, enterohemorrhagic E. coli; EIEC, enteroinvasive E. coli; EPEC, enteropathogenic E. coli; ETEC, enterotoxigenic E. coli.
sequences were then blasted against the NCBI nucleotide organism Escherichia coli, and all corresponding sequences
database to identify closely related E. coli sequences. The were selected. The sequences corresponding to each phy-
Basic Local Alignment Search Tool (BLAST; [Link] logroup genetic marker were downloaded, resulting in a
[Link]/[Link]) search parameters were set to total of five data sets. These data sets were then filtered to
yield a maximum of 5000 target sequences, maintaining retain only unique sequences to reduce processing time and
other parameters as defaults, including an expected to limit redundant results. The filtered data sets, character-
threshold of 0.05, a word size of 28, and no maximum ized by unique sequences only, were used for in silico
matches in a query range. The resulting sequences for development of probe-pairs.
each VF were compiled into a unified data set, duplicates
were removed, and ultimately, 16 data sets with unique Controls and intI1 Marker
sequence (one for each VF) were prepared for probe-pair The dMLA was designed to also include intI1, which en-
development. codes for integrase and is often associated with antibiotic-
resistant genes and horizontal gene transfer,18,30 uidA
Phylogroup Genetic Markers (used for E. coli species confirmation as a positive control),
Five phylogroup genetic markers (trpA, chuA, arpA, and gadph (used as a negative control because it is absent in
TspE4.C2, and yjaA) were selected in accordance with the E. coli, but present in Klebsiella pneumoniae). Unique
Clermont et al,29 2013, protocol to classify E. coli into reference sequences for uidA, gadph, and intI1 were ob-
corresponding phylogroups. Unique reference sequences for tained from the NCBI nucleotide database, and the same
each of these genes were obtained from the NCBI nucleo- BLAST analysis procedures were followed as described in
tide database ([Link] last accessed Phylogroup Genetic Markers ([Link]
February 17, 2021) (Supplemental Table S3). These se- gov, last accessed February 17, 2021) (Supplemental
quences were subjected to BLAST analysis against the Table S3). After filtering the results specific to each or-
NCBI nucleotide database to identify closely related se- ganism, the corresponding sequences were downloaded. For
quences, which were subsequently downloaded. The the sequences of the gadph gene, the results were filtered for
BLAST search parameters were set to yield a maximum of the organism K. pneumoniae. These data sets, now
1000 target sequences, maintaining other parameters as including intI1, uidA, and gadph, were also filtered to
defaults, including an expected threshold of 0.05, a word include only unique sequences, forming three additional
size of 28, and no maximum matches in a query range. After data sets used to develop the probe-pairs targeting the cor-
running the BLAST search, the results were filtered for the responding genes.
In Silico Probe Design Phusion buffer (1; BioLabs), 0.5 mL of dNTPs (10 nmol/
L), 1 mL of forward primer (10 mmol/L; 50 -
Probe pairs for each of the 43 genes were designed using the TCTTTTCGCAGGCTGGAGCCCAGGTCTTCCTATA-
prider package in R statistical software version 4.2.1 (http:// TAAGACTGGGCCCAATTTTCCGTGAC-3 0 ), 1 mL of
[Link]/prideR).31 Each data set of sequences, reverse primer (10 mmol/L; 50 -GCTGAACCGCTC
downloaded and generated as previously described, was TTCCGGTATCATCTTCTCCAAATGGGTCATGATC-
analysed with prider. The parameters used in prider were as 30 ), and either 1 mL of the ligation product or 1 mL of
follows: cumulative coverage decimals set to one, a mini- DNase/RNase-free water as PCR negative control.
mum primer group size of three, minimum sequence group Thermocycling was set to 98 C for 30 seconds, then 35
size of one, primer length at 40 nucleotides, and GChalves cycles of 98 C for 5 seconds, 70 C for 1 second, and
option activated. For each gene, one to five sequence clus- 72 C for 3 seconds, with a final extension at 72 C for 5
ters were identified to ensure maximum coverage of the minutes. PCR products (10 mL) were analyzed on a 2.5%
reference sequence while minimizing sequence overlap. agarose gel in TAE Buffer using 0.01% SYBR SAFE
From these clusters comprising multiple 40-bp probes that (Thermo Fisher Scientific, Schlieren, Switzerland) on a
fit the specified criteria, up to three probes per cluster were Biorad PowerPac 300 Gel Electrophoresis and multiSUB
selected for experimental validation. Each probe was then Choice Wide Midi Horizontal Electrophoresis System,
divided to generate two half-probes of 20 bp each. The 5ʹ alongside 3 mL of a 100-bp standard ladder (Promega,
end of the left half-probe was modified by adding the Madison, WI). Probe-pairs were classified as unsuc-
sequence TGGGCCCAATTTTCCGTGACAATTAATT, cessful if no band appeared at 140 bp for the positive
which incorporates the primer binding site (underlined). control or if such a band was observed in either the NTC
Similarly, the 3ʹ end of the right half-probe was extended or in the PCR negative control (Supplemental Figure S1).
with the sequence NNNNNNNNNNGAAT- If one probe-pair was classified as unsuccessful, another
GAGTGTGCGTGCACTC, which includes the reverse probe-pair from the same sequence cluster was selected
primer binding site (underlined), and NNNNNNNNNN for experimental validation. The 63 successful probe-
serves as a molecular barcode to quantify unique probe pairs targeting 43 genes are listed in Supplemental
molecules in a DNA sample and adjust for PCR amplifi- Table S4.
cation bias. Furthermore, a reverse complementary sequence
of the original 40-bp probe was used as a positive control in Multiplex on Synthetic DNA (Positive Controls)
the probe testing. Following successful single-plex performance, the 63
successful probe-pairs were selected for multiplex testing
Probe Testing and Validation of dMLA against their specific reverse cDNA templates, which
served as positive controls, alongside eight NTCs to assess
Single Plex assay specificity. Multiplex testing was performed on two
After designing the probes in silico, left and right half- technical replicates. The procedure involved mixing 1 mL
probes, along with their reverse complements serving as a of each half probe-pair from a 100 mmol/L stock to
positive control, were synthesized (Microsynth AG, Bal- generate a probe mix, which was then used in ligation
gach, Switzerland), and validated experimentally. Probe reactions. Ligation reactions comprised 1 mL of AmpLi-
pairs were tested in single plex through a ligation reaction gase (5 units; LGC Biosearch Technologies), 2.5 mL of
followed by a PCR, as previously described.21 Ligation was AmpLigase Buffer (10; LGC Biosearch Technologies),
conducted using a Biometra T3000 thermocycler (Analytik 1 mL of the probe mix to achieve a final concentration of
Jena, Jena, Germany), programmed to 94 C for 10 minutes, 1 mmol/L for each half-probe, and 1 mL of an artificial
followed by 60 C for 90 minutes, containing a total reaction DNA template (10 pmol/L) or DNase/RNase-free water for
volume of 25 mL, including the following: 1 mL of the NTCs, in a total reaction volume of 25 mL. Following
AmpLigase (5 units; Epicentre and Lucigen; LGC Bio- ligation, PCR was conducted on the products, incorpo-
search Technologies, Middleton, WI), 2.5 mL of AmpLigase rating one PCR-negative control where DNase/RNase-free
Buffer (10; Epicentre and Lucigen), 1 mL of left probe water replaced the ligation product to evaluate PCR
(1 mmol/L), 1 mL of right probe (1 mmol/L), and template. specificity. Unique barcoded forward primers for each
The template consisted of any of the following: i) 1 mL of sample (positive control DNA templates or NTCs) and a
positive control DNA template (the reverse complement uniform reverse primer were used in the PCR
sequence of both probes, 10 pmol/L), ii) 1 mL of DNase/ (Supplemental Table S5). Subsequently, 15 mL of each
RNase-free water as a ligation no template control (NTC), uniquely barcoded PCR product was pooled and purified
or iii) 1 mL of DNA extract from bacteria known to contain using the Wizard Genomic DNA Purification Kit, followed
the target gene. The subsequent PCR (25 mL total volume) by sequencing of the pooled PCR products on an Illumina
included 0.25 mL of Phusion High-Fidelity DNA Polymer- NovaSeq platform (Eurofins Genomics, GmbH, Ebersberg,
ase (2 units; New England BioLabs, Ipswich, MA), 5 mL of Germany) to generate 150-bp paired-end reads.
Application of dMLA to Bacterial Isolates lurida isolate was provided by Frederik Hammes’ laboratory
The dMLA was applied to 66 bacterial isolates, each in du- (Eawag, Duebendorf, Switzerland).
plicates, including the following: 58 E. coli, 2 Staphylococcus
aureus, 2 K. pneumoniae, 1 Klebsiella oxytoca, 1 Vibrio Sequence Data Analysis
cholera, 1 Pseudomonas lurida, and 1 Salmonella enterica.
Whole-genome sequences of the 66 bacterial isolates were Next-generation sequencing reads in FASTQ format were
available for subsequent evaluation of assay performance merged using VSEARCH version [Link] The merged
(Supplemental Table S6). Three different reactions were per- reads were processed into unique molecular barcodes counts
formed. The first reaction included 30 E. coli isolates, tested in or unique molecular identifier (UMI) counts using an ad hoc
duplicate, along with one positive control (probe-pair Python version 3.10.10 script ([Link] by
uidA160) in duplicate, eight negative controls where DNase/ matching the targets according to the forward primer
RNase-free water was used, and one PCR-negative control barcode and the target probe-pair. Data were visualized
with DNase/RNase-free water replacing the ligation product to using the package ggplot2 version 3.4.2 in R version 4.1.1
assess PCR specificity. The second reaction involved 28 E. coli via RStudio version 2024.04.2 ([Link]
isolates, also tested in duplicate, with two positive controls rstudio-desktop).
(probe-pairs aaiC47 and aggR11) in duplicate, eight negative After sequencing, UMI counts for target molecules were
controls using DNase/RNase-free water, and two PCR- observed in samples where they were not expected, indi-
negative controls to evaluate PCR specificity. The third reac- cating the presence of false positives (FPs). To minimize
tion included the eight noneE. coli isolates, tested in duplicate, these FPs, a target threshold was set. All calculations were
two positive controls (probe-pairs aaiC47 and aggR11) in performed in R version 4.1.1, with data manipulation using
duplicate, 48 negative controls using DNase/RNase-free water, the dplyr package version 1.1.1 and threshold calculation
and two PCR-negative controls to evaluate PCR specificity. performed using base R functions. The threshold was
DNA was extracted from the bacterial isolates using the Qiagen defined on the basis of the g distribution of FP UMI counts
(Hilden, Germany) Blood and Tissue kit, according to the for each target gene. Shape and rate parameters of the g
manufacturer’s instructions, starting from 1 mL of enriched distribution were estimated using the method of moments,
liquid culture in Luria Broth (AppliChem GmbH, Darmstadt, derived from the mean and variance of UMI counts identi-
Germany). The quality and concentration of the extracted DNA fied as FP. These UMI counts were determined in instances
were measured using a NanoDrop spectrophotometer (Thermo where reads corresponding to a target gene were found
Fisher Scientific), and DNA concentration was not normalized either exclusively in the presence of a positive template of a
relative to volume before using the DNA in the assay different target gene or in samples designated as negative
(Supplemental Table S7). The purification of PCR products (NTCs), thus considered FPs. Estimates were used to
and subsequent sequencing followed the same protocol as calculate the 99.9% quantile threshold for the g distribution
described for the synthetic controls, using the Wizard Genomic associated with each target. Subtraction of this threshold
DNA Purification Kit and Illumina NovaSeq platform to from the original count of each target in each sample was
generate 150-bp paired-end reads. performed, refining the detection to include only the upper
0.1% of positive reads. Any resulting negative values from
Bacterial Sample Selection this subtraction were adjusted to 0. The analyses were
Escherichia coli isolates were selected on the basis of the conducted separately for the two technical replicates and the
availability of an isolate in the laboratory with comple- three reactions performed on bacterial isolates. After
mentary whole-genome sequencing for confirmation of completing the individual analyses, the results of the three
presence/absence of gene targets, ensuring geographic and reactions on bacterial isolates were combined to evaluate the
source-based variability. Thirty isolates were collected from performance of the assay and presented in a single figure for
cattle, soil, humans, and chickens in rural villages of comprehensive visualization and comparison.
Bangladesh between February and April 2016,32 and 28
isolates were collected from Swiss wastewater in October Comparison of dMLA and Whole-Genome Sequencing
2023. NoneE. coli isolates were included to assess the Results of E. coli Isolates
specificity of the assay for E. coli, particularly targeting the
uidA gene to confirm species identity and the gapdh gene as Probe-pairs were blasted against the assembled scaffolds of
a negative control, expected to be present only in Klebsiella the whole-genome sequences of the E. coli isolates using
species. The two S. aureus strains were collected from BLAST þ version 2.15.0. The sequences of E. coli from
Swiss wastewater in December 2023 and January 2024. The set 1 were retrieved from a previous study.32 The Illumina
K. pneumoniae isolates were provided by Patrice Nord- sequencing of the isolates of E. coli from set 2 was per-
mann’s laboratory (University of Fribourg, Fribourg, formed using the LITE pipeline on a NextSeq 1000 P2 flow
Switzerland). The K. oxytoca, V. cholera, and S. enterica cell (300-bp paired end) at the Earlham Institute (Norwich,
isolates were provided by Microbial Systems Ecology UK). Sequences are available on DBJ/ENA/GenBank
Group (Eawag, Duebendorf, Switzerland), whereas the P. under the accession numbers VNWZ00000000 to
Figure 1 Heat map visualization of unique molecular identifier (UMI) counts from multiplex testing against reverse complementary templates in two
experimental replicates. The y axis displays the name of the DNA sample tested, either positive template or negative controls. The x axis indicates the name of
the probe-pairs multiplexed in the assay. Before filter: replicate 1 (A) and before filter: replicate 2 (B) show the UMI counts for each probe-pair against specific
samples before filtering with the threshold. After filter: replicate 1 (C) and after filter: replicate 2 (D) display the UMI counts after applying a 99.9% quantile
threshold of the g distribution for false-positive mitigation. Each tile corresponds to the read count of a probe-pair, with gray tiles indicating 0 UMI counts,
signifying nondetection of the gene target. The color gradient extends from green to red, representing increasing UMI counts, and turns black for UMI counts
exceeding 100,000.
samples that were expected to be positive (TP detections) pneumoniae, 1 K. oxytoca, 1 V. cholera, 1 P. lurida, and 1
had substantially higher UMI counts: 85,005 for aac3II63, S. enterica isolates (Figure 2). The assay was applied to
43,549 for aac6I123, 15,839 for dfrA369, 12,994 for each isolate in duplicate across three distinct reactions. The
gadph195, 60,893 for ipaH921, 22,605 for mcr28, and thresholds calculated as the 99.9% quantile of the g
10,835 for mcr64, indicating >800-fold increase relative to distributed FP results were applied separately to each reac-
the FPs (Supplemental Table S10). FP UMIs were rare and tion to filter out background noise in the detection data
generally showed low adjusted read counts across negative (Supplemental Table S8).
samples, with the filtering approach refined using a g model The assay successfully identified 86% of the expected
based on data from a reaction that included 50 negative positives (Figure 2). Correspondingly, 14% of the expected
controls to accurately capture background noise positives were classified as FNs. The assay correctly iden-
(Supplemental Figure S2). Finally, technical replicates dis- tified 99% of expected negative and 1% of FP (Figure 2).
played small variations in probe-pair UMI counts The expected positives were determined on the basis of the
(Supplemental Figure S3 and Supplemental Table S10). presence of the probe-pairs in the whole-genome sequences
of the isolates. In all three reactions, all NTCs and PCR-
Multiplex on Bacterial Isolates negative controls were negative (UMI counts below the
threshold), and positive controls consistently showed ex-
The multiplexed dMLA assay was conducted on 66 bacte- pected positive results in both duplicates (Figure 2 and
rial isolates, including 58 E. coli, 2 S. aureus, 2 K. Supplemental Figure S4).
Figure 2 Each cell represents the assay result for a specific probe-pair (x axis) and sample combination (y axis). The y axis displays the name of the DNA
sample tested, which includes Escherichia coli isolates (red labels), noneE. coli isolates (purple labels), synthetic DNA positive controls (blue labels), and
negative controls (green labels). The x axis indicates the name of the probe-pairs multiplexed in the assay. The color coding indicates the type of result: false
negatives (orange), false positives (yellow), true negatives (gray), and true positives (green). The expected positives were determined on the basis of the
presence of the genomic region targeted by the probe-pairs in the whole-genome sequences of the bacterial isolates. Presence of a probe-pair is indicated
when unique molecular identifier counts exceed 0 after filtering out the 99.9% quantile threshold of the g distribution for false-positive mitigation. K. oxytoca,
Klebsiella oxytoca; K. pneumoniae, Klebsiella pneumoniae; P. lurida, Pseudomonas lurida; S. aureus, Staphylococcus aureus; S. enterica, Salmonella enterica; V.
cholera, Vibrio cholera.
The assay showed high specificity across noneE. coli The assay exhibited nearly perfect specificity, achieving
isolates: the uidA marker was absent in all tested noneE. coli >99.9% in both replicates tested on synthetic DNA.
strains (Figure 2 and Supplemental Figure S4). The negative Almost all expected negatives were correctly identified as
control marker gapdh, specific to Klebsiella species, was negative, with <0.1% of the total tests resulting in FPs.
detected only in K. pneumoniae and K. oxytoca. In one of the When applied to bacterial isolates, specificity remained
two K. pneumoniae strains, six probe-pairs (aac3II63, high at 99% (Table 2).
aac6II23, aac6I6, aph6I29, intI125, and tetA13) targeting Balanced accuracy, estimated as the average of sensitivity
resistance genes were detected. Probe-pairs aph6I29 and and specificity, was high for synthetic and bacterial DNA
sul22 were also identified in S. enterica. No probe-pairs were samples, underlining the robust performance of the assay.
detected in S. aureus, P. lurida, or V. cholera isolates. Specifically, it exceeded 99.9% in both replicates tested on
Of all probe-pairs, bfpF56, bfpF105, aafA83, aap223, synthetic DNA, and reached 90% when evaluated on bac-
aggR11, eltB12, ipaH321, and qrnB35 were not detected in terial isolates (Table 2).
the whole-genome sequences of any bacterial isolate but Precision, an indication of the proportion of positives that
exhibited at least one FP result (Figure 2). Further BLAST are TPs and of negatives that are TNs, was also high for the
analysis, even with relaxed percentage identity thresholds, assay (Table 2). When tested on synthetic DNA, the preci-
confirmed the absence of these probe-pairs in the genomic sion on positive data was 95% for one replicate (indicating
sequences of all tested isolates. Conversely, probe-pairs 5% of positives are FPs) and 94% for the second replicate
dfrA15, dfrA243, dfrA369, eae6271, invE132, and qnrB7 (indicating 6% are FPs). For the bacterial isolates, precision
were not detected in any bacterial isolate, despite being was 86%, indicating that testing on bacterial isolates is
expected in specific isolates based on whole-genome associated with higher likelihood of FPs. The precision on
sequencing results, leading to FNs (Figure 2). These negative data was much higher, reaching 100% on both
probe-pairs were characterized by high threshold values for replicates of synthetic DNA and 99% on bacterial isolates.
filtering because of substantial UMI counts in the negative The F1 score, which is an indicator of balanced perfor-
controls (Supplemental Table S8), suggesting a lack of mance between sensitivity and precision, was 98% and 97%
specificity for the probe-pairs. for the two replicates tested on synthetic DNA (Table 2). As
Each bacterial isolate was tested in duplicate within the with the other indicators, the F1 score was reduced when
same reaction, showing highly similar results between du- evaluating on bacterial isolates to 83%.
plicates (Figure 2 and Supplemental Table S11). Most
samples demonstrated identical outcomes for probe detec-
Cost-Effectiveness of dMLA for Large-Scale Screening
tion in both duplicates, with most achieving 100% agree-
ment. Notably, all negative controls and positive controls The cost analysis of the dMLA for varying sample sizes in
(uidA160, aaiC47, and aggR11) also achieved 100% Switzerland, as of 2024, demonstrates low costs per isolate.
agreement, indicating consistent and reliable results across The setup costs needed to sequence one or more isolates are
both duplicates (Supplemental Table S11). However, a few approximately US$6131, including expenses for probes,
samples showed discrepancies between duplicates, where primers, enzymes, and sequencing (Table 3). For 100 iso-
certain probe-pairs were detected in one replicate but not in lates, the cost per isolate is, therefore, approximately
the other. The bacterial isolate with the largest discrepancy US$61. As much of the setup costs (ie, probes, primers) can
between duplicates nevertheless had agreement of 91% be used to characterize several thousand isolates, the costs
(n Z 57) of the probe-pairs within the multiplex, with 9% per isolate decrease as the number of isolates screened in-
(n Z 6) of the probe-pairs exhibiting different outcomes creases. For example, for 1000 isolates, the cost is reduced
(Supplemental Table S11). to US$9 per isolate; for 5000 isolates, it is further reduced to
US$5 per isolate; and for 10,000 isolates, the cost per isolate
Assay Performance Evaluation is reduced to US$4. This highlights the low costs achievable
for applying dMLA to large libraries of E. coli isolates.
The performance of the multiplex assay was evaluated
separately for each technical replicate of synthetic DNA and
for the bacterial isolates, revealing notable differences Discussion
across metrics, such as sensitivity, specificity, balanced ac-
curacy, precision, and F1 score (Table 2). The dMLA described here exhibits a high multiplexing
The assay demonstrated high sensitivity across all three capacity for targeted detection of gene regions important for
experiments. In both replicates tested on synthetic DNA, the the characterization of E. coli, including phylogroup, anti-
sensitivity was 100%, correctly identifying all expected biotic resistance, and virulence. This assay builds on the
positives as positive (Table 2). However, sensitivity promising efficacy previously reported for the digital
decreased to 81% when evaluating the assay on bacterial multiplex ligation assay method, which is designed to
isolates, with fewer anticipated gene targets accurately simultaneously screen up to hundreds of bacterial isolates.21
detected compared with synthetic templates. In this study, a total of 72 samples were screened in a single
Table 2 Predictive Performance Metrics of the Digital Multiplex Ligation Assay Across Synthetic DNA and Bacterial Isolates
Total Precision Precision
number EP EN Sensitivity, Specificity, on positive on negative F1 score, Balanced
Variable of cases cases cases TPs FPs TNs FNs % % data, % data, % % accuracy, %
Replicate 1: synthetic DNA 4536 63 4473 63 3 4470 0 100.00 99.93 95.45 97.67 99.97
Replicate 2: synthetic DNA 4536 63 4473 63 4 4471 0 100.00 99.91 94.03 96.92 99.98
Bacterial isolates 13,482 890 12,580 766 124 12,399 181 80.9 99.0 86.1 98.6 83.4 89.9
The table displays total number of possible cases, EP cases, EN cases, TPs, FPs, TNs, FNs, and the calculated values for sensitivity, specificity, precision (on
positive and on negative data), F1 score, and balanced accuracy. Sensitivity was calculated as TP/(TP þ FN), specificity as TN/(TN þ FP), precision as TP/
(TP þ FP) and TN/(TN þ FN), F1 score as 2 TP/(2 TP þ FP þ FN), and balanced accuracy as (TP/EP þ TN/EN)/2.
EN, expected negative; EP, expected positive; FN, false negative; FP, false positive; TN, true negative; TP, true positive.
run, including bacterial isolates and appropriate controls. of 2024. This is a reagents-only cost analysis and does not
The assay was further applied to an expanded suite of 19 include personnel time and equipment costs, which might
antibiotic resistance genes, 1 antimicrobial resistance vary widely across countries. Personnel and equipment costs
marker (intI1), 16 VFs, 5 E. coliespecific phylogroup for the described method are expected to be lower than for
markers, and 2 genes for species-level confirmation of E. whole-genome sequencing, given the streamlined work-
coli or K. pneumoniae. The assay demonstrated high accu- flows and fewer technical requirements. For comparison, the
racy and performance when tested on both synthetic DNA cost of whole-genome sequencing in different countries
and a library of DNA extracts of bacterial isolates, where the around the world ranges from US$72 to US$470 per isolate,
uidA marker or other unintended targets were not detected in as published in a recent review.37 Moreover, further cost
noneE. coli isolates, and gapdh was detected only in reductions could be achieved by increasing the number of
Klebsiella species, as expected. False positives or misclas- barcoded primers (namely, by pooling together a higher
sification can occur if samples do not represent individual number of samples for sequencing, with the primary limi-
bacterial isolates. For example, if isolation of bacterial tation theoretically driven only by sequencing depth).21 In
cultures leads to enrichment of more than one bacterial low-and middle-income countries, where screening for
isolate (eg, a co-culture of E. coli and Klebsiella species). pathogenic and antimicrobial resistant markers in bacterial
The dMLA is a highly promising method for high- isolates is often limited, the cost-efficiency and high-
throughput screening, allowing simultaneous detection of throughput nature of dMLA can significantly improve the
multiple gene regions and the screening of hundreds of monitoring of pathogenic and antimicrobial resistant E. coli
bacterial isolates in a single run.21 The ability to pool PCR in environmental samples.38
products from multiple samples, each marked with a unique Implementing dMLA to screen large E. coli collections
barcode, significantly reduces sequencing costs, making it could offer a comprehensive view of prevalent pathogenic
an efficient and cost-effective method for extensive sample and resistant strains, supporting more targeted and region-
analysis. The sequencing costs to detect the presence of 43 ally effective public health interventions. Such an approach
genes could be reduced to approximately US$4 to US$5 per could be particularly beneficial in strengthening programs
sample when scaling up to screen 5000 to 10,000 isolates, as like the Multiple Indicator Cluster Survey, which assesses
Table 3 Cost Analysis of Digital Multiplex Ligation Assay for Different Sample Sizes
Maximum isolates per Initial costs per 100 Isolates, 1000 Isolates, 5000 Isolates, 10,000 Isolates,
Variable stock (estimated) stock (estimated), US$ US$ US$ US$ US$
Left probes 50,000 575 575 575 575 575
Right probes 50,000 3106 3106 3106 3106 3106
Reverse primer 5000 4 4 4 4 8
Forward primer 5000 1140 1140 1140 1140 2280
Ligase enzyme 600 369 369 738 3321 6273
DNA polymerase 2000 664 664 664 1992 3320
Sequencing 100 274 274 2740 13,700 27,400
(5 million read pairs)
Total costs e 6131 6131 8966 23,837 42,961
Costs/isolate e e 61 9 5 4
The table shows the initial setup costs and the resulting cost per isolate for 100, 1000, 5000, and 10,000 samples. The table includes the maximum number
of samples that can be processed with one stock of each component. The higher cost of the right probe compared with the left probe is due to 50 phos-
phorylation required for ligation. Forward primers are more expensive than reverse primers because they include molecular barcodes for multiplexing. The
sequencing cost of US$274 per 100 samples reflects the pooled sequencing of purified PCR products from all samples.
e, Not applicable.
drinking water quality using E. coli as an indicator of fecal occurrences, aligning with the explanation of thresholding-
contamination in low-and middle-income countries, but driven variability in UMI counts. A higher number of
does not differentiate between pathogenic and resistant negative samples in each assay might provide a more ac-
strains.22 curate distribution to estimate the threshold for detect/non-
Although the assay only achieved 90% balanced accuracy detect classification. Ensuring a reliable distribution helps in
when tested on bacterial isolates, its primary utility is setting thresholds that minimize the inclusion of FPs. Non-
attributed to the high-throughput screening capability. specific amplification and probe binding to nontarget se-
Lower balanced accuracy indicates that there is a disparity quences can also contribute to FPs.40 Although primer-
in the detection of TPs and TNs. Specifically, it was found dimer formation was assessed in silico, such interactions
that the assay on bacterial isolates has a higher rate of FNs, may not always be detectable computationally. In a complex
for specific probe-pairs more than for others, and it also real matrix, chemical reactions during ligation may still
occasionally produced FPs by incorrectly identifying nega- allow for non-specific binding or cross-reactivity, such as
tive samples as positive. probes binding to other probes. Similarly, carry-over
False negatives might be explained by high UMI counts contamination or other inefficiencies in the assay process
observed in the negative controls, which caused the esti- may cause the observed low concentration of target reads in
mated thresholds (used to filter reads to differentiate be- all samples independent from the true presence of the target
tween a detect and nondetect) to also be high. Specifically, gene. If the number of these reads is low in DNA samples,
when negative samples exhibit high UMI counts for a even if the number is low in the negative controls, the
probe-pair, the threshold for detecting a TP for the corre- threshold might not be high enough to filter them out,
sponding probe-pair will be high. If this threshold is equal to resulting in low-count FPs. Additionally, laboratory prac-
or exceeds the UMI counts in TP samples, it can result in tices, such as cross-contamination during sample process-
misclassifying these TPs as (falsely) negatives. High UMI ing, can also lead to FPs.40 Similarly, FPs might also be
counts in negative samples could be due to non-specific attributed to background noise introduced during the library
binding or cross-reactivity, such as probes binding to other preparation for sequencing.
probes during ligationdan affinity not detected in silico. A The dMLA assay demonstrated high specificity, with a g
higher rate of FNs was observed among probe-pairs tar- distribution model applied to FP UMIs to accurately set
geting different regions of the same gene, suggesting detection thresholds. Adjusting these thresholds further risks
possible competition during the ligation process or cross- misclassifying TPs, particularly for probe-pairs with low
binding in negative samples, which may compromise UMI counts in positive samples. Notably, the reliability of
assay performance by reducing the limit of detection and the filtering threshold increases with the inclusion of a
increasing the probability of misrepresentative results.39 higher number of negative samples, which provides a more
When specific primer-probes are problematic for FNs (eg, accurate estimate of the background distribution.
the dfrA369 probes), redesigning to target different gene The FPs reported in this study may also result from
regions may help to reduce specificity issues in nondetects, imperfect reference genomes to determine the ground truth.
or otherwise increase assay performance. In this study, FPs may have resulted from inaccurate non-
False positives were observed as well, which could be detection of the gene in the whole-genome sequences. For
attributed to several factors. False positives with low UMI instance, genes located on plasmids may not be detected if
counts can be influenced by the distribution of UMI counts the plasmid is lost during the culturing process or during
in the negative samples. The cause of the detection of a DNA extraction, or if the plasmid DNA is underrepresented
target gene in only one NTC in the study (with an adjusted in the sequencing data because of its lower abundance
UMI count of one for mcr64) is unknown, but it is likely compared with chromosomal DNA.41 Additionally, plasmid
due to contamination during sample preparation. All NTCs sequences can be more difficult to assemble and identify in
in each replicate were prepared using the same nuclease-free whole-genome sequencing data, leading to their absence in
water; if the water was contaminated, it likely would have the assembled genome but their presence being detected by
affected many or all of the NTCs. Additionally, presence of the assay.41 Incomplete or fragmented genome assemblies
the target gene before PCR amplification would have where certain regions are missing or poorly covered can also
resulted in significantly higher post-filtration UMI counts, contribute to the nondetection of these genes in the
resembling UMI counts observed in TPs with synthetic sequencing data.42
DNA controls. Other potential sources of the FP cannot be Despite these challenges, the potential to multiplex
excluded, such as cross-contamination of amplified products probes for comprehensive characterization of bacterial iso-
after PCR with low levels during or after sequencing steps, lates using dMLA is high, offering a robust first step for
or primer-dimer formation. However, this appears to be a high-throughput screening. dMLA may also offer an initial
stochastic process, as no FPs were detected in the NTCs screening approach, which can be followed by PCR
during bacterial isolate testing. The fact that only 1 of 16 confirmation for enhanced reliability or for helping to
NTCs across two replicates on synthetic DNA showed an identify isolates for subsequent whole-genome sequencing.
FP underlines the random and infrequent nature of such Indeed, in scenarios where extensive sample analysis is
4. Cosgrove SE: The relationship between antimicrobial resistance and 21. Tamminen M, Spaak J, Caduff L, Schiff H, Lang R, Schmid S,
patient outcomes: mortality, length of hospital stay, and health care Montealegre MC, Julian TR: Digital multiplex ligation assay for
costs. Clin Infect Dis 2006, 42:S82eS89. Suppl 2 highly multiplexed screening of [beta]-lactamase-encoding genes in
5. World Health Organization: Global action plan on antimicrobial bacterial isolates. Commun Biol 2020, 3:1e6
resistance. Geneva, Switzerland: WHO, 2015. Available at: https:// 22. Khan S, Hancioglu A: Multiple indicator cluster surveys: delivering
[Link]/publications/i/item/9789241509763. Accessed March robust data on children and women across the globe. Stud Fam Plann
13, 2025 2019, 50:279e286
6. Seale AC, Gordon NC, Islam J, Peacock SJ, Scott JAG: AMR sur- 23. Miller WR, Arias CA: ESKAPE pathogens: antimicrobial resistance,
veillance in low and middle-income settingsda roadmap for partic- epidemiology, clinical impact and therapeutics. Nat Rev Microbiol
ipation in the Global Antimicrobial Surveillance System (GLASS). 2024, 2024:1e19
Wellcome Open Res 2017, 2:92 24. Zhang AN, Gaston JM, Dai CL, Zhao S, Poyet M, Groussin M,
7. World Health Organization: Antimicrobial resistance: global report Yin X, Li LG, van Loosdrecht MCM, Topp E, Gillings MR,
on surveillance. Geneva, Switzerland: WHO, 2014. Available at: Hanage WP, Tiedje JM, Moniz K, Alm EJ, Zhang T: An omics-based
[Link] Accessed framework for assessing the health risk of antimicrobial resistance
April 17, 2025 genes. Nat Commun 2021, 12:4765
8. Diallo OO, Baron SA, Abat C, Colson P, Chaudet H, Rolain JM: 25. Jahantigh M, Samadi K, Dizaji RE, Salari S: Antimicrobial resistance
Antibiotic resistance surveillance systems: a review. J Glob Anti- and prevalence of tetracycline resistance genes in Escherichia coli
microb Resist 2020, 23:430e438 isolated from lesions of colibacillosis in broiler chickens in Sistan,
9. World Health Organization: Global antimicrobial resistance and use Iran. BMC Vet Res 2020, 16:267
surveillance system (GLASS) report: 2022. Geneva, Switzerland: 26. Wang Z, Fu Y, Schwarz S, Yin W, Walsh TR, Zhou Y, He J, Jiang H,
WHO, 2022. Available at: [Link] Wang Y, Wang S: Genetic environment of colistin resistance genes
9789240062702. Accessed April 17, 2025 mcr-1 and mcr-3 in Escherichia coli from one pig farm in China. Vet
10. World Health Organization: WHO integrated global surveillance on Microbiol 2019, 230:56e61
ESBL-producing E. coli using a “One Health” approach: imple- 27. Jiang H, Cheng H, Liang Y, Yu S, Yu T, Fang J, Zhu C: Diverse
mentation and opportunities. Geneva, Switzerland: WHO, 2021. mobile genetic elements and conjugal transferability of sulfonamide
Available at: [Link] resistance genes (sul1, sul2, and sul3) in Escherichia coli isolates
402. Accessed April 17, 2025 from Penaeus vannamei and pork from large markets in Zhejiang,
11. European Centre for Disease Prevention and Control and World China. Front Microbiol 2019, 10:1787
Health Organization Regional Office for Europe: Antimicrobial 28. Chen L, Yang J, Yu J, Yao Z, Sun L, Shen Y, Jin Q: VFDB: a
resistance surveillance in Europe 2023d2021 data. Copenhagen, reference database for bacterial virulence factors. Nucleic Acids Res
Denmark: WHO, 2023. Available at: [Link] 2005, 33:D325eD328
publications/i/item/9789289058537. Accessed April 17, 2025 29. Clermont O, Christenson JK, Denamur E, Gordon DM: The Clermont
12. World Health Organization: Central Asian and European surveillance Escherichia coli phylo-typing method revisited: improvement of
of antimicrobial resistance: annual report 2019. Geneva, Switzerland: specificity and detection of new phylo-groups. Environ Microbiol
WHO, 2019. Available at: [Link] Rep 2013, 5:58e65
item/9789289058179. Accessed April 17, 2025 30. Baltazar M, Bourgeois-Nicolaos N, Larroudé M, Couet W,
13. Ashley EA, Recht J, Chua A, Dance D, Dhorda M, Thomas NV, Uwajeneza S, Doucet-Populaire F, Ploy MC, Da Re S: Activation of
Ranganathan N, Turner P, Guerin PJ, White NJ, Day NP: An in- class 1 integron integrase is promoted in the intestinal environment.
ventory of supranational antimicrobial resistance surveillance net- PLoS Genet 2022, 18:e1010177
works involving low- and middle-income countries since 2000. J 31. Smolander N, Julian TR, Tamminen M: Prider: multiplexed primer
Antimicrob Chemother 2018, 73:1737e1749 design using linearly scaling approximation of set coverage. BMC
14. Buckley GJ, Palmer GH: Combating Antimicrobial Resistance and Bioinformatics 2022, 23:174
Protecting the Miracle of Modern Medicine. Washington, DC, Na- 32. Montealegre MC, Talavera Rodríguez A, Roy S, Hossain MI,
tional Academies Press, 2022 Islam MA, Lanza VF, Julian TR: High genomic diversity and
15. Schreiber C, Zacharias N, Essert SM, Wasser F, Müller H, Sib E, heterogenous origins of pathogenic and antibiotic-resistant Escher-
Precht T, Parcina M, Bierbaum G, Schmithausen RM, Kistemann T, ichia coli in household settings represent a challenge to reducing
Exner M: Clinically relevant antibiotic-resistant bacteria in aquatic transmission in low-income settings. mSphere 2020, 5:1e17
environments e an optimized culture-based approach. Sci Total En- 33. Rognes T, Flouri T, Nichols B, Quince C, Mahé F: VSEARCH: a
viron 2021, 750:142265 versatile open source tool for metagenomics. PeerJ 2016, 4:e2584
16. Laprade N, Cloutier M, Lapen DR, Topp E, Wilkes G, Villemur R, 34. Powers DMW: Evaluation: from precision, recall and F-measure to
Khan IUH: Detection of virulence, antibiotic resistance and toxin ROC, informedness, markedness and correlation. arXiv 2020, [Pre-
(VAT) genes in Campylobacter species using newly developed print], [Link]
multiplex PCR assays. J Microbiol Methods 2016, 124:41e47 35. Taha AA, Hanbury A: Metrics for evaluating 3D medical image
17. Keenum I, Liguori K, Calarco J, Davis BC, Milligan E, Harwood VJ, segmentation: analysis, selection, and tool. BMC Med Imaging 2015,
Pruden A: A framework for standardized qPCRdtargets and protocols 15:1e28
for quantifying antibiotic resistance in surface water, recycled water and 36. Mower JP: PREP-Mt: predictive RNA editor for plant mitochondrial
wastewater. Crit Rev Environ Sci Technol 2022, 52:4395e4419 genes. BMC Bioinformatics 2005, 6:96
18. Manaia CM: Framework for establishing regulatory guidelines to 37. Price V, Ngwira LG, Lewis JM, Baker KS, Peacock SJ, Jauneikaite E,
control antibiotic resistance in treated effluents. Crit Rev Environ Sci Feasey N: A systematic review of economic evaluations of whole-
Technol 2023, 53:754e779 genome sequencing for the surveillance of bacterial pathogens.
19. Guo J, Li J, Chen H, Bond PL, Yuan Z: Metagenomic analysis re- Microb Genom 2023, 9:000947
veals wastewater treatment plants as hotspots of antibiotic resistance 38. Iskandar K, Molinier L, Hallit S, Sartelli M, Hardcastle TC,
genes and mobile genetic elements. Water Res 2017, 123:468e478 Haque M, Lugova H, Dhingra S, Sharma P, Islam S, Mohammed I,
20. Ferreira C, Otani S, Aarestrup FM, Manaia CM: Quantitative PCR Naina Mohamed I, Hanna PA, Hajj SEl, Jamaluddin NAH,
versus metagenomics for monitoring antibiotic resistance genes: Salameh P, Roques C: Surveillance of antimicrobial resistance in low-
balancing high sensitivity and broad coverage. FEMS Microbes 2023, and middle-income countries: a scattered picture. Antimicrob Resist
4:xtad008 Infect Control 2021, 10:63
39. Tighe PJ, Ryder RR, Todd I, Fairclough LC: ELISA in the multiplex 43. Browne DJ, Kelly AM, Brady JL, Doolan DL: A high-throughput
era: potentials and pitfalls. Proteomics Clin Appl 2015, 9:406 screening RT-qPCR assay for quantifying surrogate markers of im-
40. Borst A, Box ATA, Fluit AC: False-positive results and contamina- munity from PBMCs. Front Immunol 2022, 13:962220
tion in nucleic acid amplification assays: suggestions for a prevent 44. Dreier M, Meola M, Berthoud H, Shani N, Wechsler D, Junier P:
and destroy strategy. Eur J Clin Microbiol Infect Dis 2004, 23: High-throughput qPCR and 16S rRNA gene amplicon sequencing as
289e299 complementary methods for the investigation of the cheese micro-
41. Arredondo-Alonso S, Willems RJ, van Schaik W, Schürch AC: biota. BMC Microbiol 2022, 22:1e18
On the (im)possibility of reconstructing plasmids from whole- 45. Hedges DJ, Guettouche T, Yang S, Bademci G, Diaz A, Andersen A,
genome short-read sequencing data. Microb Genom 2017, 3: Hulme WF, Linker S, Mehta A, Edwards YJK, Beecham GW,
e000128 Martin ER, Pericak-Vance MA, Zuchner S, Vance JM, Gilbert JR:
42. Peona V, Blom MPK, Xu L, Burri R, Sullivan S, Bunikis I, Liachko I, Comparison of three targeted enrichment strategies on the SOLiD
Haryoko T, Jønsson KA, Zhou Q, Irestedt M, Suh A: Identifying the sequencing platform. PLoS One 2011, 6:e18595
causes and consequences of assembly gaps using a multiplatform 46. Quince C, Walker AW, Simpson JT, Loman NJ, Segata N: Shotgun
genome assembly of a bird-of-paradise. Mol Ecol Resour 2021, 21: metagenomics, from sampling to analysis. Nat Biotechnol 2017, 35:
263 833e844