Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species
Abstract
Background - The process of generating raw genome sequence data continues to become cheaper, faster, and more accurate. However, assembly of such data into high-quality, finished genome sequences remains challenging. Many genome assembly tools are available, but they differ greatly in terms of their performance (speed, scalability, hardware requirements, acceptance of newer read technologies) and in their final output (composition of assembled sequence). More importantly, it remains largely unclear how to best assess the quality of assembled genome sequences. The Assemblathon competitions are intended to assess current state-of-the-art methods in genome assembly. Results - In Assemblathon 2, we provided a variety of sequence data to be assembled for three vertebrate species (a bird, a fish, and snake). This resulted in a total of 43 submitted assemblies from 21 participating teams. We evaluated these assemblies using a combination of optical map data, Fosmid sequences, and several statistical methods. From over 100 different metrics, we chose ten key measures by which to assess the overall quality of the assemblies. Conclusions - Many current genome assemblers produced useful assemblies, containing a significant representation of their genes, regulatory sequences, and overall genome structure. However, the high degree of variability between the entries suggests that there is still much room for improvement in the field of genome assembly and that approaches which work well in assembling the genome of one species may not necessarily work well for another.
Keywords
Cite
@article{arxiv.1301.5406,
title = {Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species},
author = {Keith R. Bradnam and Joseph N. Fass and Anton Alexandrov and Paul Baranay and Michael Bechner and İnanç Birol and Sébastien Boisvert and Jarrod A. Chapman and Guillaume Chapuis and Rayan Chikhi and Hamidreza Chitsaz and Wen-Chi Chou and Jacques Corbeil and Cristian Del Fabbro and T. Roderick Docking and Richard Durbin and Dent Earl and Scott Emrich and Pavel Fedotov and Nuno A. Fonseca and Ganeshkumar Ganapathy and Richard A. Gibbs and Sante Gnerre and Élénie Godzaridis and Steve Goldstein and Matthias Haimel and Giles Hall and David Haussler and Joseph B. Hiatt and Isaac Y. Ho and Jason Howard and Martin Hunt and Shaun D. Jackman and David B Jaffe and Erich Jarvis and Huaiyang Jiang and Sergey Kazakov and Paul J. Kersey and Jacob O. Kitzman and James R. Knight and Sergey Koren and Tak-Wah Lam and Dominique Lavenier and François Laviolette and Yingrui Li and Zhenyu Li and Binghang Liu and Yue Liu and Ruibang Luo and Iain MacCallum and Matthew D MacManes and Nicolas Maillet and Sergey Melnikov and Bruno Miguel Vieira and Delphine Naquin and Zemin Ning and Thomas D. Otto and Benedict Paten and Octávio S. Paulo and Adam M. Phillippy and Francisco Pina-Martins and Michael Place and Dariusz Przybylski and Xiang Qin and Carson Qu and Filipe J Ribeiro and Stephen Richards and Daniel S. Rokhsar and J. Graham Ruby and Simone Scalabrin and Michael C. Schatz and David C. Schwartz and Alexey Sergushichev and Ted Sharpe and Timothy I. Shaw and Jay Shendure and Yujian Shi and Jared T. Simpson and Henry Song and Fedor Tsarev and Francesco Vezzi and Riccardo Vicedomini and Jun Wang and Kim C. Worley and Shuangye Yin and Siu-Ming Yiu and Jianying Yuan and Guojie Zhang and Hao Zhang and Shiguo Zhou and Ian F. Korf},
journal= {arXiv preprint arXiv:1301.5406},
year = {2015}
}
Comments
Additional files available at http://korflab.ucdavis.edu/Datasets/Assemblathon/Assemblathon2/Additional_files/ Major changes 1. Accessions for the 3 read data sets have now been included 2. New file: spreadsheet containing details of all Study, Sample, Run, & Experiment identifiers 3. Made miscellaneous changes to address reviewers comments. DOIs added to GigaDB datasets