Discovery of tandem and interspersed segmental duplications using high-throughput sequencing

Söylev, Arda; Le, T. M.; Amini, H.; Alkan, Can; Hormozdiari, F.

Discovery of tandem and interspersed segmental duplications using high-throughput sequencing

buir.contributor.author	Söylev, Arda
buir.contributor.author	Alkan, Can
dc.citation.epage	3930	en_US
dc.citation.issueNumber	20	en_US
dc.citation.spage	3923	en_US
dc.citation.volumeNumber	35	en_US
dc.contributor.author	Söylev, Arda	en_US
dc.contributor.author	Le, T. M.	en_US
dc.contributor.author	Amini, H.	en_US
dc.contributor.author	Alkan, Can	en_US
dc.contributor.author	Hormozdiari, F.	en_US
dc.date.accessioned	2020-02-12T10:58:57Z
dc.date.available	2020-02-12T10:58:57Z
dc.date.issued	2019-04
dc.department	Department of Computer Engineering	en_US
dc.description.abstract	Motivation: Several algorithms have been developed that use high-throughput sequencing technology to characterize structural variations (SVs). Most of the existing approaches focus on detecting relatively simple types of SVs such as insertions, deletions and short inversions. In fact, complex SVs are of crucial importance and several have been associated with genomic disorders. To better understand the contribution of complex SVs to human disease, we need new algorithms to accurately discover and genotype such variants. Additionally, due to similar sequencing signatures, inverted duplications or gene conversion events that include inverted segmental duplications are often characterized as simple inversions, likewise, duplications and gene conversions in direct orientation may be called as simple deletions. Therefore, there is still a need for accurate algorithms to fully characterize complex SVs and thus improve calling accuracy of more simple variants. Results: We developed novel algorithms to accurately characterize tandem, direct and inverted interspersed segmental duplications using short read whole genome sequencing datasets. We integrated these methods to our TARDIS tool, which is now capable of detecting various types of SVs using multiple sequence signatures such as read pair, read depth and split read. We evaluated the prediction performance of our algorithms through several experiments using both simulated and real datasets. In the simulation experiments, using a 30 coverage TARDIS achieved 96% sensitivity with only 4% false discovery rate. For experiments that involve real data, we used two haploid genomes (CHM1 and CHM13) and one human genome (NA12878) from the Illumina Platinum Genomes set. Comparison of our results with orthogonal PacBio call sets from the same genomes revealed higher accuracy for TARDIS than state-of-the-art methods. Furthermore, we showed a surprisingly low false discovery rate of our approach for discovery of tandem, direct and inverted interspersed segmental duplications prediction on CHM1(<5% for the top 50 predictions).	en_US
dc.identifier.doi	10.1093/bioinformatics/btz237	en_US
dc.identifier.eissn	1460-2059	en_US
dc.identifier.issn	1367-4803	en_US
dc.identifier.uri	http://hdl.handle.net/11693/53302	en_US
dc.language.iso	English	en_US
dc.publisher	Oxford University Press	en_US
dc.relation.isversionof	https://dx.doi.org/10.1093/bioinformatics/btz237	en_US
dc.source.title	Bioinformatics	en_US
dc.title	Discovery of tandem and interspersed segmental duplications using high-throughput sequencing	en_US
dc.type	Article	en_US

Files

Original bundle

Now showing 1 - 1 of 1

Name:: Discovery_of_tandem_and_interspersed_segmental_duplications_using_high-throughput_sequencing.pdf
Size:: 678.34 KB
Format:: Adobe Portable Document Format
Description:

Download

Collections

Scholarly Publications - Computer Engineering