Rapid alignment-free phylogenetic identification of metagenomic sequences

Benjamin Linard 1 Krister Swenson 2 Fabio Pardi 1
1 MAB - Méthodes et Algorithmes pour la Bioinformatique
LIRMM - Laboratoire d'Informatique de Robotique et de Microélectronique de Montpellier
Abstract : Motivation: Taxonomic classification is at the core of environmental DNA analysis. When a phyloge-netic tree can be built as a prior hypothesis to such classification, phylogenetic placement (PP) provides the most informative type of classification because each query sequence is assigned to its putative origin in the tree. This is useful whenever precision is sought (e.g. in diagnostics). However, likelihood-based PP algorithms struggle to scale with the ever-increasing throughput of DNA se-quencing.
Results: We have developed RAPPAS (Rapid Alignment-free Phylogenetic Placement via Ancestral Sequences) which uses an alignment-free approach, removing the hurdle of query sequence alignment as a preliminary step to PP. Our approach relies on the precomputation of a database of k-mers that may be present with non-negligible probability in relatives of the reference sequences. The placement is performed by inspecting the stored phylogenetic origins of the k-mers in the query, and their probabilities. The database can be reused for the analysis of several different metagenomes. Experiments show that the first implementation of RAPPAS is already faster than competing likelihood based PP algorithms, while keeping similar accuracy for short reads. RAPPAS scales PP for the era of routine metagenomic diagnostics.
Availability: Program and sources freely available for download at https://github.com/blinard-BIOINFO/RAPPAS
Complete list of metadatas

https://hal.archives-ouvertes.fr/hal-02008297
Contributor : Benjamin Linard <>
Submitted on : Tuesday, February 5, 2019 - 4:06:31 PM
Last modification on : Thursday, May 9, 2019 - 10:47:43 AM
Long-term archiving on : Monday, May 6, 2019 - 4:12:23 PM

File

btz068.pdf
Publisher files allowed on an open archive

Licence


Distributed under a Creative Commons Attribution - NonCommercial 4.0 International License

Identifiers

Collections

Citation

Benjamin Linard, Krister Swenson, Fabio Pardi. Rapid alignment-free phylogenetic identification of metagenomic sequences. Bioinformatics, Oxford University Press (OUP), 2019, ⟨10.1093/bioinformatics/btz068/5303992⟩. ⟨hal-02008297⟩

Share

Metrics

Record views

56

Files downloads

22