13 papers · ranked by Valyu relevance
Raelynn Chen, Attri Ghosh, Jie Hu, Yong Chen + 2 more
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does…
Raelynn Chen, Attri Ghosh, Jie Hu, Yong Chen + 2 more
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does…
Nazila Ahmadi Daryakenari, Seyed Kamaleddin Setaredan
Schizophrenia (SZ) is a chronic and complex mental disorder associated with neurobiological deficits. The complexity and heterogeneity of schizophrenia symptoms pose challenges for objective diagnosis, which is currently based on behavioral and clinical manifestations. Furthermore, other psychiatric disorders such as…
Sina Kanannejad, Noemi Bongiorni, Elisa Nordera, Sara Redaelli + 4 more
Single-cell RNA sequencing has enabled the construction of comprehensive cell atlases, yet the quality and coherence of the cell-type annotations within these atlases remain largely unexamined. When a label is applied to a transcriptionally heterogeneous population, the downstream analyses that depend on it, and…
Erik D. VonKaenel, Lisa M. Bramer, Javier E. Flores, Thomas O Metz + 2 more
In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by…
Philipp S. L. Schäfer, Kendall A. Reid, Zach Boldyga, Ekin D. Aksu + 2 more
Single-cell perturbation experiments measure how interventions alter cellular phenotypes. However, the number of possible perturbations and biological contexts far exceeds what can be tested experimentally. Motivated by this constraint, predictive models aim to extrapolate cellular responses to unseen conditions.…
Marcelo Hurtado, Vera Pancaldi
Machine learning approaches are increasingly applied to high-dimensional biological data in which features are often dataset-dependent. In many omics workflows, features are computed using information derived from the entire dataset, such as correlations between variables, clustering structures, or enrichment scores.…
Shiri Baum, Ido Meshulam, Yadid M Algavi, Omri Peleg + 1 more
TAGINE is a feature engineering algorithm that leverages the microbial taxonomic tree to optimize feature sets in microbiome data for predictive modeling. The algorithm starts with features at high taxonomic levels and iteratively splits them into lower-level clades in cases where it improves predictive accuracy…
Yang Zhang, Lin Liu, Lina Ma, Zhang Zhang
Multi-omics integrative analysis is pivotal for elucidating complex molecular mechanisms and biological processes, yet remains challenging in multi-omics data integration and feature selection. Here we present MIA, a machine learning framework for multi-omics integrative analysis that features unsupervised sample…
Jin-Zhao He, Jing Guan
Identifying prognostic biomarkers from high-dimensional transcriptomic data poses a triple challenge: achieving sparsity, preserving biological network topology, and integrating complementary nonlinear signals. Existing methods typically ignore network structure, miss nonlinear interactions, or lack a principled…
Rossana O. Souza, Wellington Francisco Rodrigues, Bráulio R. G. M. Couto, Marcos A. dos Santos
Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic…
Haibin Guan, Maaike van Gerwen, Seunghee Kim-Schultz, Elena Colicino + 2 more
High-dimensional multi-omics data integration offers novel opportunities to characterize complex biological systems. Even though sampling variability frequently compromises findings, particularly in small cohorts, the reproducibility and generalizability of the derived latent structures are insufficiently evaluated. We…
Yijiang Liu, Yuting Wang, Tao Huan, Xiaotao Shen
Liquid chromatography-mass spectrometry (LC-MS) untargeted metabolomics detects thousands of metabolic features, but converting these chemical signals into metabolite set-level biological knowledge remains challenging. This is because most features lack unambiguous metabolite identities. Conventional metabolite set…