14 papers · ranked by Valyu relevance
Raelynn Chen, Attri Ghosh, Jie Hu, Yong Chen + 2 more
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does…
Raelynn Chen, Attri Ghosh, Jie Hu, Yong Chen + 2 more
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does…
Erik D. VonKaenel, Lisa M. Bramer, Javier E. Flores, Thomas O Metz + 2 more
In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by…
Cheyenne N. Jarman, Taal Levi, Mark Novak
Applications of machine learning in ecology are rapidly expanding. Symbolic regression is gaining particular attention for its success in reverse-engineering human-readable explanatory population models, including the logistic growth and Lotka-Volterra equations, from simulated and laboratory-based population time…
Zhikang Liu, Yiyang Niu, Tian Le, Daniel G Chen + 2 more
The rapid maturation of single-cell multi-omics technologies has enabled unprecedented resolution for mapping disease states and identifying disease-associated biomarkers. In practice, biomarkers are often discovered through differential detection that treat genomic features as independent contributors to phenotypes…
Rebecca Danning, Zheng Tracy Ke, Xihong Lin, Rong Ma
The prioritization of highly-variable genes is an important step in single-cell trajectory inference. However, when variability arises from a continuous latent cell development trajectory, standard methods may fail to differentiate trajectory-relevant from uninformative genes. SEEK-VFI is an ensemble topic-modeling…
Jinwoo Lee, Junghoon Justin Park, Maria Pak, Seung Yun Choi + 1 more
Analyzing individual differences in treatment or exposure effects is a central challenge in psychology and behavioral sciences. Conventional statistical models have focused on average treatment effects, overlooking individual variability, and struggling to identify key moderators. Generalized Random Forest (GRF) can…
Adham M. Alkhadrawi, Mohammed A.B. Mahmoud, Mian M.Y. Khalil, Abdullah All Jaber
Feature selection is a critical preprocessing step in single-cell RNA sequencing (scRNA-seq) analysis, directly impacting downstream clustering and biological interpretation. We systematically compared 16 feature selection methods across three diverse datasets: PBMC3K (immune cells), Visium Heart, and Visium Brain…
Fang Nan, David Azriel, Armin Schwartzman
High-dimensional genetic data present substantial challenges for estimating the fraction of variance explained (FVE) by genome-wide single-nucleotide polymorphisms (SNPs). In the context of genetics the VFE is called SNP heritability. Standard approaches for FVE estimation, such as GWAS heritability (GWASH) and linkage…
Rossana O. Souza, Wellington Francisco Rodrigues, Bráulio R. G. M. Couto, Marcos A. dos Santos
Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic…
Alexander Henoch, Metehan Sever, Sarah J. Tucker, Florian Trigodet + 7 more
Pangenomics quantifies the conserved and variable gene repertoire among genomes, but popular implementations ignore gene synteny. Graph-based approaches incorporate both gene homology and synteny, but become difficult to interpret due to pervasive rearrangements. Here we present network-pruning and graph-layout…
Kengo Sakurai, Laurence Moreau, Tristan Mary-Huard, Alain Charcosset + 1 more
In plant breeding, it is often necessary to improve a target trait while maintaining other essential traits within desirable ranges. When genetic relationships exist among these traits, improvements in the target trait may lead to undesirable changes in essential traits, complicating cross selections. In such cases, it…
Burak Yelmen, Merve Nur Güler, Tõnu Kollo, Märt Möls + 2 more
Over the past two decades, genome-wide association studies (GWAS) enabled the discovery of thousands of variants associated with many complex human traits. However, conventional GWAS are still widely performed with linear models with the assumption that the genetic effects are predominantly additive. In this work, we…
Ruoxuan Wu, Xiudi Li, Feifei Xiao, Muxuan Liang
Mendelian randomization (MR), leveraging genetic variants as instrumental variables (IVs), is widely used to draw causal conclusions in the presence of unmeasured confounding, but most MR analyses focus on average treatment effects and rely on strong assumptions. For precision medicine, the primary target is instead…