15 papers · ranked by Valyu relevance
Marika Ström, Nicole Wagner, Iryna Kolosenko, Åsa M. Wheelock
The R-workflow ropls-ViPerSNet (R orthogonal projections of latent structures with Variable Permutation Selection and Elastic Net) facilitates variable selection, model optimization and significance testing using permutations of OPLS-DA models, with the scaled loadings (p[corr]) as the main metric of significance…
Wei Cheng, Sohini Ramachandran, Lorin Crawford
In this paper, we propose a new approach for variable selection using a collection of Bayesian neural networks with a focus on quantifying uncertainty over which variables are selected. Motivated by fine-mapping applications in statistical genetics, we refer to our framework as an “ensemble of single-effect neural…
Roberta Coletti, J. Orestes Cerdeira, Marcos Raydan, Marta B. Lopes
High-dimensional omics data often contain more variables than observations, which negatively impacts the performance of classical data analysis methods. Dimensionality reduction is typically addressed through variable selection strategies that incorporate a penalty term into the model. While effective for selecting…
Raelynn Chen, Attri Ghosh, Jie Hu, Yong Chen + 2 more
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does…
Guannan Yang, Ellen Menkhorst, Evdokia Dimitriadis, Kim-Anh Lê Cao
The knockoff framework, combined with variable selection procedure, controls false discovery rate (FDR) without the need for calculating p−values. Hence, it presents an attractive alternative to differential expression analysis of high-throughput biological data. However, current knockoff variable generators make…
Xiangyu Zhang, Lijun Wang, Jia Zhao, Hongyu Zhao
Transcriptome-wide association studies (TWASs) have been developed to nominate candidate genes associated with complex traits by integrating genome-wide association studies (GWASs) with expression quantitative trait loci (eQTL) data. However, most existing TWAS methods evaluate the marginal association between a single…
Raelynn Chen, Attri Ghosh, Jie Hu, Yong Chen + 2 more
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does…
Tingting Zhao, Guangyu Zhu, Patrick Flaherty
Large-scale multiple perturbation experiments have the potential to reveal a more detailed understanding of the molecular pathways that respond to genetic and environmental changes. A key question in these studies is which gene expression changes are important for the response to the perturbation. We present here a…
Robert Dunne
Random Forests (RF) are a very widely used modelling tool. 34 concludes that no nonlinear model had a more widespread popularity, from health care to academia to industry, than random forests and decision trees. The bounds of the methodology are still being extended. 4 give an example with 80 million variables. It is…
Anirban Samaddar, Tapabrata Maiti, Gustavo de los Campos
Variable selection and large-scale hypothesis testing are techniques commonly used to analyze high-dimensional genomic data. Despite recent advances in theory and methodology, variable selection and inference with highly collinear features remain challenging. For instance, collinearity poses a great challenge in…
Yusuke Imoto
Accurate selection of highly variable genes (HVGs) is essential in single-cell RNA sequencing (scRNA-seq) data analysis, as it enables the identification of functionally important genes and the characterization of cell types and states. However, HVG selection is often confounded by technical noise inherent in the…
Erik D. VonKaenel, Lisa M. Bramer, Javier E. Flores, Thomas O Metz + 2 more
In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by…
Suruchi Jai Kumar Ahuja
A major objective of clustering is to identify groups in the data that maximizes the similarity between objects within the same cluster and minimizes the similarity between different clusters. A challenge for data clustering, and unsupervised learning in general, is that there is often no mechanism for feature…
Marija Kekic, Oleg Stepanov, Wenjuan Wang, Sam Richardson + 7 more
Covariate selection in population pharmacokinetics modelling is essential for understanding interindividual variability in drug response and optimizing dosing. Traditional stepwise covariate modelling is often time-consuming, compared to the new machine learning alternatives. This study investigates the use of Neural…
Stijn Hawinkel, Olivier Thas, Steven Maere
The winner’s curse is a form of selection bias that arises when estimates are obtained for a large number of features, but only a subset of most extreme estimates is reported. It occurs in large scale significance testing as well as in rank-based selection, and imperils reproducibility of findings and follow-up study…