Data, time and money: evaluating the best compromise for inferring molecular phylogenies of non-model animal taxa

Zaharias, Paul; Pante, Eric; Gey, Delphine; Fedosov, Alexander e.; Puillandre, Nicolas

doi:10.1016/j.ympev.2019.106660

Data, time and money: evaluating the best compromise for inferring molecular phylogenies of non-model animal taxa

Type

Article

Date

2020-01

Author(s)

Zaharias Paul ¹, Pante Eric ², Gey Delphine ³, Fedosov Alexander e. ⁴, Puillandre Nicolas ¹

Affiliation(s)

Univ Antilles, Sorbonne Univ, Inst Systemat Evolut Biodiversite ISYEB, Museum Natl Hist Nat,CNRS,EPHE, 43 Rue Cuvier,CP 26, F-75005 Paris, France.
La Rochelle Univ, UMR 7266 CNRS, Littoral Environm & Soc LIENSs, 2 Rue Olympe de Gouges, F-17042 La Rochelle, France.
Museum Natl Hist Nat, Acquisit & Anal Donnees Hist Nat 2AD UMS 2700, Paris, France.
Russian Acad Sci, AN Severtsov Inst Ecol & Evolut, Leninsky Prospect 33, Moscow 119071, Russia.

Source

Molecular Phylogenetics And Evolution (1055-7903) (Academic Press Inc Elsevier Science), 2020-01, Vol. 142, P. 106660 (11p.)

DOI

10.1016/j.ympev.2019.106660

Archimer ID

86802

WOS© Times Cited

18

Language(s)

English

Publisher

Academic Press Inc Elsevier Science

For over a decade now, High Throughput sequencing (HTS) approaches have revolutionized phylogenetics, both in terms of data production and methodology. While transcriptomes and (reduced) genomes are increasingly used, generating and analyzing HTS datasets remain expensive, time consuming and complex for most nonmodel taxa. Indeed, a literature survey revealed that 74% of the molecular phylogenetics trees published in 2018 are based on data obtained through Sanger sequencing. In this context, our goal was to identify the strategy that would represent the best compromise among costs, time and robustness of the resulting tree. We sequenced and assembled 32 transcriptomes of the marine mollusk family Turridae, considered as a typical non-model animal taxon. From these data, we extracted the loci most commonly used in gastropod phylogenies (cox1, 12S, 16S, 28S, h3 and 18S), full mitogenomes, and a reduced nuclear transcriptome representation. With each dataset, we reconstructed phylogenies and compared their robustness and accuracy. We discuss the impact of missing data and the use of statistical tests, tree metrics, and supertree and supermatrix methods to further improve phylogenetic data acquisition pipelines. We evaluated the overall costs (time and money) in order to identify the best compromise for phylogenetic data sampling in non-model animal taxa. Although sequencing full mitogenomes seems to constitute the best compromise both in terms of costs and node support, they are known to induce biases in phylogenetic reconstructions. Rather, we recommend to systematically include loci commonly used for phylogenetics and taxonomy (i.e. DNA barcodes, rRNA genes, full mitogenomes, etc.) among the other loci when designing baits for capture.

Keyword(s)

Phylogenomics, Transcriptomics, High throughput sequencing, Sanger sequencing, Non-model taxa, Turridae

Full Text

File	Pages	Size
Publisher's official version	11	1 Mo
Author's final draft	49	1 Mo	Download
Supplementary Fig. 1. 20 species tree produced for this study.	20	391 Ko
Supplementary Fig. 2. Distribution of quartet distance of single-locus trees of the UPh-AS16 dataset against the UPh-AS16 supertree.	1	11 Ko
Supplementary Table 1. Description of the specimens and transcriptomes.	-	14 Ko
Supplementary Table 2. Correlation table between different sequencing and assembly results.	-	8 Ko
Supplementary Table 3. Quartet scores for ASTRAL-III datasets.	-	7 Ko
Supplementary Table 4. Evaluation of the costs (time and money) for each dataset.	-	15 Ko
Supplementary Table 5. Correlation coefficient of single-loci’s quartet distance against several alignment statistics.	-	7 Ko

How to cite

Zaharias Paul, Pante Eric, Gey Delphine, Fedosov Alexander E., Puillandre Nicolas (2020). Data, time and money: evaluating the best compromise for inferring molecular phylogenies of non-model animal taxa. Molecular Phylogenetics And Evolution. 142. 106660 (11p.). https://doi.org/10.1016/j.ympev.2019.106660, https://archimer.ifremer.fr/doc/00756/86802/

Copy this text