24 papers · ranked by Valyu relevance
Ramón Díaz-Uriarte, Sara Alvarez de Andrés
Background Selection of relevant genes for sample classification is a common task in most gene expression studies, where researchers try to identify the smallest possible set of genes that can still achieve good predictive performance (for instance, for future use with diagnostic purposes in clinical practice). Many…
Yusuke Imoto
Accurate selection of highly variable genes (HVGs) is essential in single-cell RNA sequencing (scRNA-seq) data analysis, as it enables the identification of functionally important genes and the characterization of cell types and states. However, HVG selection is often confounded by technical noise inherent in the…
Justin Bleich, Adam Kapelner, Edward I. George, Shane T. Jensen
We consider the task of discovering gene regulatory networks, which are defined as sets of genes and the corresponding transcription factors which regulate their expression levels. This can be viewed as a variable selection problem, potentially with high dimensionality. Variable selection is especially challenging in…
Muhammad Hamraz, Naz Gul, Mushtaq Raza, Dost Muhammad Khan + 4 more
'Umair Khalil' 'Seema Zubair' 'Zardad Khan' 'Ka-Chun Wong'] In this paper, a novel feature selection method called Robust Proportional Overlapping Score (RPOS), for microarray gene expression datasets has been proposed, by utilizing the robust measure of dispersion, i.e., Median Absolute Deviation (MAD). This method…
Gao Wang, Abhishek Sarkar, Peter Carbonetto, Matthew Stephens
We introduce a simple new approach to variable selection in linear regression, and to quantifying uncertainty in selected variables. The approach is based on a new model – the “Sum of Single Effects” (SuSiE) model – which comes from writing the sparse vector of regression coefficients as a sum of “single-effect”…
Shuzhen Sun, Zhuqi Miao, Blaise Ratcliffe, Polly Campbell + 4 more
High-throughput sequencing technology has revolutionized both medical and biological research by generating exceedingly large numbers of genetic variants. The resulting datasets share a number of common characteristics that might lead to poor generalization capacity. Concerns include noise accumulated due to the large…
Carmen Lai, Marcel JT Reinders, Laura J van't Veer, Lodewyk FA Wessels
'Lodewyk FA Wessels'] Background Gene selection is an important step when building predictors of disease state based on gene expression data. Gene selection generally improves performance and identifies a relevant subset of genes. Many univariate and multivariate gene selection approaches have been proposed. Frequently…
Emmanuel P. Dollinger, Kai Silkwood, Scott Atwood, Qing Nie + 1 more
Background The high dimensionality of data in single cell transcriptomics (scRNAseq) requires investigators to choose subsets of genes (“feature selection”) for downstream analysis (e.g., unsupervised cell clustering). The evaluation of different approaches to feature selection is hampered by the fact that, as we show…
Shuzhen Sun, Zhuqi Miao, Blaise Ratcliffe, Polly Campbell + 5 more
With the rapid advancement of DNA sequencing technology, the volume and dimension of biological and medical data have been increasing at an unprecedented rate. Accompanying such high volume genetic data, the ‘curse of dimensionality’ has challenged the validity of statistical methods that do not scale to massive data.…
Souvik Bag, Kapil Gupta, Soudeep Deb
The selection of essential variables in logistic regression is vital because of its extensive use in medical studies, finance, economics and related fields. In this paper, we explore four main typologies (test-based, penalty-based, screening-based, and tree-based) of frequentist variable selection methods in logistic…
David Kepplinger, Peter Filzmoser, Курт Вармуза
Genetic algorithms are a widely used method in chemometrics for extracting variable subsets with high prediction power. Most fitness measures used by these genetic algorithms are based on the ordinary least-squares fit of the resulting model to the entire data or a subset thereof. Due to multicollinearity, partial…
Perrine Lacroix, Mélina Gallopin, Marie‐Laure Martin‐Magniette
Variable selection methods are widely used in molecular biology to detect biomarkers or to infer gene regulatory networks from transcriptomic data. Methods are mainly based on the highdimensional Gaussian linear regression model and we focus on this framework for this review. We propose a comparison study of variable…
Zhenqiang Su, Huixiao Hong, Hong Fang, Leming Shi + 2 more
'Weida Tong'] Background Advances in DNA microarray technology portend that molecular signatures from which microarray will eventually be used in clinical environments and personalized medicine. Derivation of biomarkers is a large step beyond hypothesis generation and imposes considerably more stringency for accuracy…
Chee Chun Gan, Gerard P. Learmonth
Genetic algorithms are a well-known method for tackling the problem of variable selection. As they are non-parametric and can use a large variety of fitness functions, they are well-suited as a variable selection wrapper that can be applied to many different models. In almost all cases, the chromosome formulation used…
Subrata Saha, Ahmed Soliman, Sanguthevar Rajasekaran
Nowadays we are observing an explosion of gene expression data with phenotypes. It enables researchers to efficiently identify genes responsible for certain medical condition as well as classify them for drug target. Like any other phenotype data in medical domain, gene expression data with phenotypes also suffers from…
Nilotpal Sanyal
High-dimensional data with binary outcomes are common in biological sciences and related fields such as omics, epidemiology, and healthcare, environmental sciences, physical and engineering sciences, and business and policy making. Often, such data are sparse so that only a small number of all available variables truly…
Damir Zhakparov, Kathleen Moriarty, Damian Roqueiro, Katja Baerenfaller
High-dimensional Bulk RNA sequencing (RNAseq) datasets pose a considerable challenge in identifying biologically relevant features for downstream analyses and data mining efforts. The standard approach involves differential gene expression (DGE) analysis, but its effectiveness can be limited depending on the data due…
Long Ma, Nancy J. Lin, Christopher I. Amos, Momiao Xiong
Fast and more economical next generation sequencing (NGS) technologies will generate unprecedentedly massive and highly-dimensional genomic and epigenomic variation data. In the near future, a routine part of medical records will include the sequenced genomes. How to efficiently extract biomarkers for risk prediction…
Michael Lecocke, Kenneth Hess
Background We consider both univariate- and multivariate-based feature selection for the problem of binary classification with microarray data. The idea is to determine whether the more sophisticated multivariate approach leads to better misclassification error rates because of the potential to consider jointly…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Georgia A. Henry, John R. Stinchcombe
Evolution by natural selection occurs at its most basic through the change in frequencies of alleles; connecting those genomic targets to phenotypic selection is an important goal for evolutionary biology in the genomics era. The relative abundance of gene products expressed in a tissue can be considered a phenotype…
Authors not listed
Ensuring the trustworthiness of machine learning (ML) models in high-stake applications is crucial. One such application is predicting anti-cancer drug sensitivity, where ML models are built with the final goal of integrating them into treatment recommendation systems for personalized medicine. Here, we propose a…
Authors not listed
One aim of the international Human Proteome Organization (HUPO) Human Proteome Project (HPP) is to obtain high-confidence translation evidence for every human protein-coding gene established in its target list of 19433 entries based on the protein-coding genes from Ensembl-GENCODE. However, 76 are annotated in…
Authors not listed
This paper addresses the challenges in cell line development (CLD), the lengthy and ambiguous clone screening in upstream biopharmaceutical production. Typically, only a small subset of the later stages of CLD data is used for manually selecting lead clones. Addressing this issue, we introduce a multivariate data…