Taxonomic tools




















However, not all methods improved during the second run, indicating that their databases were not cache friendly Figure 7A. The disparity in runtime between a cached and uncached database can be up to x, so careful use of caching and batch processing can result in major performance improvements.

Advancements in the field of metagenomic sequencing analysis have produced a suite of taxonomic classifiers that perform well across a range of datasets, meaning that other preferences or constraints should also guide method selection. The classifiers investigated here had similar performances, especially between classifiers of the same type, on the same dataset. Broadly, DNA classifiers provide better precision-recall and abundance estimates than protein-based classifiers when using a uniform database for these whole-genome datasets.

This lower specificity of protein classifiers on the normalized RefSeq CG database can be largely attributed to the absence of non-coding sequence from their databases. Indeed, differences in default database compositions accounted for greater performance differences between classifiers than the methods themselves.

This similarity in performance means that other factors such as the computational resources available, as well as the ease-of-use and other particulars of each classifier should be considered when deciding on the most appropriate tool. The computational resource requirements of some programs may be prohibitive for some applications, so using classifiers that require lower memory incurs a tradeoff between computational cost and accuracy.

If custom databases are desired, Centrifuge may be useful; it requires 10s of Gbs of memory and demonstrates good performance metrics, despite the shortcomings of the compressed default database. Among the protein-based classifiers, Kaiju is generally recommended for having much faster classification speed and lower memory requirements, compared to the other classifiers DIAMOND and MMseqs2 , without compromising performance.

All three of these methods allow custom databases. Although most of the classifiers included in this study performed well across the benchmark datasets containing known taxa, there are several areas in which more development is needed to improve these tools. This includes both simple technical changes to existing methods as well as conceptual innovations in how these methods are used that could improve the quality of the results generated by these classifiers. Classifiers perform very well when taxa in a sample are genetically distinct from each other and genetically similar to sequences in the reference database.

Though not evaluated here, previous studies have shown that classification below the species level is challenging with current taxonomic classifiers McIntyre et al. The dramatically poorer performance of all classifiers on the CAMI datasets compared to the IMMSA datasets, further underlines the influence of evolutionary distance and poorly described taxa on classification performance.

Relative classifier performance was largely consistent, indicating that these problems are common across existing tools. Expanding reference databases can improve classification Forster et al. De novo assembly of metagenomes may be a more appropriate strategy for analysis of such highly complex and novel metagenomic samples.

These tools have recently identified thousands of novel metagenome assembled genomes in the human microbiome Almeida et al. One of the biggest performance challenges for many classifiers is that they often report large numbers of low abundance false positives, lowering the precision of these estimates. Most practitioners filter these reports using a given abundance threshold. Integration of this within software packages would improve classifier precision and standardisation and simplify downstream analyses.

Another possibility for improving performance would be performing per-read filtering based on alignment score or e-value. Some classifiers have options that allow for this, but do not have default recommendations for using them. These values would depend on classifier support and would vary between classifiers and even experimental designs, which could make them difficult to use.

When the host is known, pre-filtering of host-derived sequences can also be used Blauwkamp et al. More generally, making full use of information available about sequenced reads - including incorporating sequence quality scores and PCR duplicate removal - common practices in other areas of sequence analysis, could further improve classification. Using quality scores in a probabilistic context would increase the computational cost of classification, but could be a fruitful avenue for future research.

Similarly, most classifiers generate profiles simply based on proportions of classified reads. Of the methods evaluated here, only Bracken takes a probabilistic approach to generate the final abundance profiles.

Another common issue in low-input metagenomic sequencing is high PCR duplication rates of sequencing reads resulting from high numbers of PCR cycles. This can distort abundance estimates since certain taxa become dominant via high duplicate numbers. Some effort has been made by methods like KrakenUniq to estimate the unique nucleotide content of each taxa, to better determine duplication level and recover true abundances.

In addition to PCR duplication, there are a number of other features of the experimental preparation of samples for metagenomic sequencing that can confound classification performance. Cross-talk between multiplexed samples on Illumina flow-cells especially notable on more recent sequencers that use patterned flow cells can generate biological false positives at low abundances Sinha et al. Library molecules can bleed through multiple runs on the sample sequencers.

Samples can also pick up biological contamination from microbial sequences in the laboratory environment or genetic material present in extraction and library preparation reagents.

These contaminants may confound benchmarking attempts using real sequenced datasets, however, this is part of a larger challenge of metagenomic classification that needs to be addressed. Clean laboratory practices as well as recent innovations to experimentally remove routine contaminants Gu et al. The use of well-matched negative controls in any metagenomics study is essential to be able to identify and control for contamination.

Though use of negative controls has recently improved in metagenomics studies, there is currently no accepted approach for correcting for contamination using these controls Davis et al. One interesting avenue for exploration is the use of negative control samples to create blacklists of known contaminants. Thinking more broadly, one proposed strategy to address the individual shortcomings of different classifiers is to use an ensemble of classifiers, or a pipeline of different classifiers McIntyre et al.

For pipelines, typically a faster method is run first, while a slower but more sensitive method is run on the leftover unclassified or poorly classified reads Bazinet et al. This is sometimes performed in a more ad hoc fashion using protein classifiers as a second-pass analysis option for DNA classifiers because it offers higher sensitivity to mutations and more distantly related proteins, but trades lower specificity in response Yang et al.

The downside of ensemble classifiers is higher computational runtime and a more challenging interpretation of the results. Runtime is a recurring challenge for many of these methods and though some methods use approaches that either reduce runtime or memory use, many classifiers have runtimes that are dictated by the disk speed of reading their databases into memory.

After loading the database, classifying each incremental sample is extremely fast for methods such as Kraken and CLARK. Not all methods support classifying multiple samples in a single execution to explicitly take advantage of database caching, yet their database pages might by implicitly cached by the operating system for faster single-sample executions.

These methods are faster with larger amounts of memory and fast temporary disks. The rapid growth of reference databases will present a fundamental challenge to the field in coming years as this pace exceeds that of computer memories. This will likely create a number of secondary challenges that current software will need to adjust to.

The growth of reference databases such as RefSeq can result in methods changing performance characteristics over time, as exemplified by the decrease in accuracy of k-mer based methods Nasko et al. Many of these newly sequenced genomes are re-sequenced microbial strains that are very similar to existing species, resulting in oversampling of some genera and species groups. There are some ways that software programs can mitigate taxonomy changes, such as including the full versioned taxonomy database tree for hosted pre-built metagenomic databases.

However, these changes do not address the underlying challenge presented by the size and growth of these databases. Managing this data explosion is likely to be one of the biggest challenges in metagenomic sequence classification in the coming years and will require more wide-reaching changes in our approach. The field of metagenomics is approaching a critical milestone in its trajectory. Investment in recent years has led to the development of a range of programs with good overall performance, giving users a choice of ways to analyse their data in accordance with their particular question, computational environment, target taxa and other preferences.

This is making the analysis of metagenomic sequencing more accessible than ever before. However, many taxonomic classifiers are still burdened by high numbers of false positive calls at low abundance that need to be addressed.

Looking beyond this, bigger breakthroughs will be needed in many areas, including in controlling experimental sources of contamination and error, and in handling the exponential growth of reference databases, in order to create transformational changes in metagenomic classification towards microbial detection and characterization.

This project was also funded in part by a Broadnext10 gift from the Broad Institute, and by the Bill and Melinda Gates foundation. Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. As a service to our customers we are providing this early version of the manuscript. The manuscript will undergo copyediting, typesetting, and review of the resulting proof before it is published in its final citable form.

Please note that during the production process errors may be discovered which could affect the content, and all legal disclaimers that apply to the journal pertain.

The authors declare no competing interests. National Center for Biotechnology Information , U. Author manuscript; available in PMC Aug 8. Author information Copyright and License information Disclaimer. Author Contributions Conceptualization S. Methodology, Software, Formal Analysis, S. Writing - Original Draft, K. S and S. Writing - Review and Editing, D.

Y Supervision, D. Copyright notice. The publisher's final edited version of this article is available at Cell. See other articles in PMC that cite the published article. Associated Data Supplementary Materials 1. Summary Metagenomic sequencing is revolutionizing the detection and characterization of microbial species, and a wide variety of software tools are available to perform taxonomic classification of this data. Open in a separate window. Figure 1. Efficient classification of millions of reads A large number of tools have recently been developed that are focused on classifying large amounts of sequencing reads to known taxa with increasing speed.

Size and growth of reference databases All metagenomics classifiers require a pre-computed database based on previously sequenced microbial genetic sequences, whose sheer size presents a considerable computational challenge.

Comparing classifier performance The metrics selected to benchmark classifiers can greatly influence their relative rankings and performance and thus must be carefully selected to best reflect the way these tools are used in practice. Figure 2. Evaluating the precision-recall of 20 classifiers Here we benchmarked 20 metagenomic classifiers to compare performance in classification precision, recall, F1, speed, and other metrics using a uniform database to eliminate any confounding effects of differences in default databases.

Table 1. Figure 3. Comparing abundance profile distances for 20 classifiers We next evaluated the accuracy of the estimated abundance profiles across classifiers compared to the ground truth. Figure 4. Figure 5. Determining the rate and source of false positive classifications False positive classifications present a major challenge for the interpretation of metagenomic sequencing data, especially when considering human clinical samples White et al.

Figure 6. Figure 7. Discussion Advancements in the field of metagenomic sequencing analysis have produced a suite of taxonomic classifiers that perform well across a range of datasets, meaning that other preferences or constraints should also guide method selection.

Supplementary Material 1 Click here to view. Footnotes Publisher's Disclaimer: This is a PDF file of an unedited manuscript that has been accepted for publication. Nucleic Acids Res. Analyzing and minimizing PCR amplification bias in Illumina sequencing libraries. Genome Biol. The Statistical Analysis of Compositional Data. Series B Stat. A new genomic blueprint of the human gut microbiota. Nature , — Binning metagenomic contigs by coverage and composition.

Methods 11 , — Basic local alignment search tool. Normalization methods for microbial abundance data strongly affect correlation estimates. BLAST-based validation of metagenomic sequence assignments. PeerJ 6 , e Analytical and clinical validation of a microbial cell-free DNA sequencing test for infectious disease.

Nat Microbiol 4 , — KrakenUniq: confident and fast metagenomics classification using unique k-mer counts. Methods 12 , 59— Clinical metagenomics. Genome Res. A comprehensive benchmarking study of protocols and sequencing platforms for 16S rRNA community profiling.

BMC Genomics 17 , Simple statistical identification and removal of contaminant sequences in marker-gene and metagenomics data. Microbiome 6 , Bioinformatics 34 , — Opportunistic Data Structures with Applications. A human gut bacterial genome and culture collection for improved metagenomic analyses. Depletion of Abundant Sequences by Hybridization DASH : using Cas9 to remove unwanted high-abundance species in sequencing libraries and molecular counting applications.

Plant Sci. MetaBAT, an efficient tool for accurately reconstructing single genomes from complex microbial communities. PeerJ 3 , e Centrifuge: rapid and sensitive classification of metagenomic sequences. Bayesian community-wide culture-independent microbial source tracking. If you continue browsing the site, you agree to the use of cookies on this website. See our User Agreement and Privacy Policy. See our Privacy Policy and User Agreement for details. Create your free account to read unlimited documents.

The SlideShare family just got bigger. Home Explore Login Signup. Successfully reported this slideshow. We use your LinkedIn profile and activity data to personalize ads and to show you more relevant ads. You can change your ad preferences anytime. Next SlideShares. You are reading a preview. Create your free account to continue reading. Sign Up. Upcoming SlideShare. What to Upload to SlideShare. Embed Size px. Start on.

Marine Planktonic Ostracods : M. Ostracods from waters around Great Britain. Includes extensive references, figures, maps of all described species of copepods. Kasturirangan, Aetideidae of the World Ocean : by E. Also available on dvd. Covers all described species with an image, description, references, etc. A collection of classic taxonomic texts: free downloads : copepods, medusae, ctenophores. Giesbrecht, W. Plankton of the offshore waters of the Gulf of Maine.

By Henry B. Available on the Biodiversity Heritage Library website. Rose, M. Faune de France, ed. The World of Copepods : Bibliography, taxonomic list, databases of researchers, specimens, genera. Also, links to other copepod sites.



0コメント

  • 1000 / 1000