15 papers · ranked by Valyu relevance
Anthony J. Greenberg
Explosive growth in the amount of genomic data is matched by increasing power of consumer-grade computers. Even applications that require powerful servers can be quickly tested on desktop or laptop machines if we can generate representative samples from large data sets. I describe a fast and memory-efficient…
Anirduddha Laud, Gaurav Menghani, Madhava Keralapura
However, human genomes are highly redundant. Any given individual's genome would differ from another individual's genome by less than 1%. There are tools like DNAZip, which express a given genome sequence by only noting down the differences between the given sequence and a reference genome sequence. This allows…
Zhanshan, Lianwei Li, Ya‐Ping Zhang
—Classic concepts of genetic (gene) diversity (heterozygosity) such as Nei (1973: PNAS) and Nei & Li (1979: PNAS) nucleotide diversity were defined within the context of populations. Although variations are often measured in population context, the basic carriers of variation are individuals. Hence, measuring…
Travis Gagie
Indexing large genomic databases is a challenging problem in bioinformatics, and several authors have designed heavily-engineered data structures specifically for this purpose; see, e.g., [1] and references therein. In this paper we propose a simple model of these databases and give a theoretical solution, which we…
Carlos W. Nossa, Paul Havlak, Jia‐Xing Yue, Jie Lv + 3 more
'Kimberly Y Vincent' 'H. Jane Brockmann' 'Nicholas H. Putnam'] - Carlos Nossa1 , Paul Havlak1 , Jia-Xing Yue1 , Jie Lv1 , Kim Vincent1 , H Jane Brockmann3 4 , Nicholas H Putnam1,2 6 (1) Department of Ecology and Evolutionary Biology, and (2) Department of Biochemistry and Cell Biology, Rice University, P.O. Box 1892…
Fathima Nuzla Ismail, Shanika Amarasoma
—Next-generation sequencing (NGS) is a pivotal technique in genome sequencing due to its high throughput, rapid results, cost-effectiveness, and enhanced accuracy. Its significance extends across various domains, playing a crucial role in identifying genetic variations and exploring genomic complexity. NGS finds…
Jaroslav Budiš, Werner Krampl, Marcel Kucharík, Rastislav Hekel + 11 more
'Adrian Goga' 'Michal Lichvár' 'Dávid Smoľak' 'Miroslav Böhmer' 'Andrej Baláž' 'František Ďuriš' 'Juraj Gazdarica' 'Katarína Šoltýs' 'Ján Turňa' 'Ján Radvánszky' 'Tomáš Szemes'] Geneton Ltd., 841 04 Bratislava, Slovakia Slovak Centre of Scientific and Technical Information, 811 04 Bratislava, Slovakia Comenius…
Chang‐Yong Lee
Motivated by a non-random but clustered distribution of SNPs, we introduce a phenomenological model to account for the clustering properties of SNPs in the human genome. The phenomenological model is based on a preferential mutation to the closer proximity of existing SNPs. With the Hapmap SNP data, we empirically…
Umadevi Paila, Brad Chapman, Rory Kirchner, Aaron R. Quinlan
Modern DNA sequencing technologies enable geneticists to rapidly identify genetic variation among many human genomes. However, isolating the minority of variants underlying disease remains an important, yet formidable challenge for medical genetics. We have developed GEMINI (GEnome MINIng), a flexible software package…
James Lindesay, Tshela E. Mason, William Hercules, Georgia M. Dunston
'Georgia M. Dunston'] - 1 Computational Physics Laboratory, Howard University, Washington, DC, 20059, U.S. E-mail: jlindesay@howard.edu (J.L.) - 2 National Human Genome Center, Howard University, Washington, DC, 20060, U.S. E-mail: tmason@howard.edu (T.E.M.); gdunston@howard.edu (G.M.D.) - 3 Department of Physics…
Saulo Aflitos, Gabino Sanchez‐Perez, Dick de Ridder, Paul Fransz + 3 more
'M. Eric Schranz' 'Hans de Jong' 'Sander Peters'] Breeding by introgressive hybridization is a pivotal strategy to broaden the genetic basis of crops. Usually, the desired traits are monitored in consecutive crossing generations by marker-assisted selection, but their analyses fail in chromosome regions where crossover…
Günter Jäger, Alexander Peltzer, Kay Nieselt
Background: To understand individual genomes it is necessary to look at the variations that lead to changes in phenotype and possibly to disease. However, genotype information alone is often not sufficient and additional knowledge regarding the phase of the variation is needed to make correct interpretations.…
Andrew Oldfield, William Ritchie
Results: We developed SeqManager, a web-based application that provides automated identification, classification, and management of sequencing data files with intelligent duplicate detection. It also detects intermediate sequencing files that can safely be removed. Evaluation across four genomics laboratory settings…
Anunchai Assawamakin, Nachol Chaiyaratana, Chanin Limwongse, Saravudh Sinsomros + 2 more
'Saravudh Sinsomros' 'Pa‐thai Yenchitsomanus' 'Prakarnkiat Youngkong'] Abstract— This paper presents a non-parametric classification technique for identifying a candidate bi-allelic genetic marker set that best describes disease susceptibility in gene-gene interaction studies. The developed technique functions by…
Swetansu Pattnaik, Saurabh Gupta, Arjun A. Rao, Binay Panda
We report SInC (SNV, Indel and CNV) simulator and read generator, an open-source tool capable of simulating biological variants taking into account a platform-specific error model. SInC is capable of simulating and generating single- and paired-end reads with user-defined insert size with high efficiency compared to…