Aggressive assembly of pyrosequencing reads with mates
- PMID: 18952627
- PMCID: PMC2639302
- DOI: 10.1093/bioinformatics/btn548
Aggressive assembly of pyrosequencing reads with mates
Abstract
Motivation: DNA sequence reads from Sanger and pyrosequencing platforms differ in cost, accuracy, typical coverage, average read length and the variety of available paired-end protocols. Both read types can complement one another in a 'hybrid' approach to whole-genome shotgun sequencing projects, but assembly software must be modified to accommodate their different characteristics. This is true even of pyrosequencing mated and unmated read combinations. Without special modifications, assemblers tuned for homogeneous sequence data may perform poorly on hybrid data.
Results: Celera Assembler was modified for combinations of ABI 3730 and 454 FLX reads. The revised pipeline called CABOG (Celera Assembler with the Best Overlap Graph) is robust to homopolymer run length uncertainty, high read coverage and heterogeneous read lengths. In tests on four genomes, it generated the longest contigs among all assemblers tested. It exploited the mate constraints provided by paired-end reads from either platform to build larger contigs and scaffolds, which were validated by comparison to a finished reference sequence. A low rate of contig mis-assembly was detected in some CABOG assemblies, but this was reduced in the presence of sufficient mate pair data.
Availability: The software is freely available as open-source from http://wgs-assembler.sf.net under the GNU Public License.
Figures
Similar articles
-
An algorithm for automated closure during assembly.BMC Bioinformatics. 2010 Sep 10;11:457. doi: 10.1186/1471-2105-11-457. BMC Bioinformatics. 2010. PMID: 20831800 Free PMC article.
-
QuorUM: An Error Corrector for Illumina Reads.PLoS One. 2015 Jun 17;10(6):e0130821. doi: 10.1371/journal.pone.0130821. eCollection 2015. PLoS One. 2015. PMID: 26083032 Free PMC article.
-
Benchmarking of de novo assembly algorithms for Nanopore data reveals optimal performance of OLC approaches.BMC Genomics. 2016 Aug 22;17 Suppl 7(Suppl 7):507. doi: 10.1186/s12864-016-2895-8. BMC Genomics. 2016. PMID: 27556636 Free PMC article.
-
Assembly algorithms for next-generation sequencing data.Genomics. 2010 Jun;95(6):315-27. doi: 10.1016/j.ygeno.2010.03.001. Epub 2010 Mar 6. Genomics. 2010. PMID: 20211242 Free PMC article. Review.
-
Sequence assembly using next generation sequencing data--challenges and solutions.Sci China Life Sci. 2014 Nov;57(11):1140-8. doi: 10.1007/s11427-014-4752-9. Epub 2014 Oct 17. Sci China Life Sci. 2014. PMID: 25326069 Review.
Cited by
-
Molecular epidemiologic investigation of an anthrax outbreak among heroin users, Europe.Emerg Infect Dis. 2012 Aug;18(8):1307-13. doi: 10.3201/eid1808.111343. Emerg Infect Dis. 2012. PMID: 22840345 Free PMC article.
-
Unbiased approach for virus detection in skin lesions.PLoS One. 2013 Jun 28;8(6):e65953. doi: 10.1371/journal.pone.0065953. Print 2013. PLoS One. 2013. PMID: 23840382 Free PMC article.
-
Denoising DNA deep sequencing data-high-throughput sequencing errors and their correction.Brief Bioinform. 2016 Jan;17(1):154-79. doi: 10.1093/bib/bbv029. Epub 2015 May 29. Brief Bioinform. 2016. PMID: 26026159 Free PMC article.
-
Genome sequence of the model medicinal mushroom Ganoderma lucidum.Nat Commun. 2012 Jun 26;3:913. doi: 10.1038/ncomms1923. Nat Commun. 2012. PMID: 22735441 Free PMC article.
-
HASLR: Fast Hybrid Assembly of Long Reads.iScience. 2020 Aug 21;23(8):101389. doi: 10.1016/j.isci.2020.101389. Epub 2020 Jul 25. iScience. 2020. PMID: 32781410 Free PMC article.
References
-
- Bentley DR. Whole-genome re-sequencing. Curr. Opin. Genet. Dev. 2006;16:545–552. - PubMed
-
- Blattner FR, et al. The complete genome sequence of Escherichia coli K-12. Science. 1997;277:1453–1474. - PubMed
-
- Chou HH, Holmes MH. DNA sequence quality trimming and vector removal. Bioinformatics. 2001;17:1093–1104. - PubMed
-
- Denisov G, et al. Consensus generation and variant detection by Celera Assembler. Bioinformatics. 2008;24:1035–1040. - PubMed
Publication types
MeSH terms
Grants and funding
LinkOut - more resources
Full Text Sources
Other Literature Sources