23 papers · ranked by Valyu relevance
Jon Donnelly, Srikar Katta, Emanuele Borgonovo, Cynthia Rudin
Variable importance (VI) methods are often used for hypothesis generation, feature selection, and scientific validation. In the standard VI pipeline, an analyst estimates VI for a single predictive model with only the observed features. However, the importance of a feature depends heavily on which other variables are…
Paweł Morzywołek, Peter B. Gilbert, Alex Luedtke
We provide an inferential framework to assess variable importance for heterogeneous treatment effects. This assessment is especially useful in high-risk domains such as medicine, where decision makers hesitate to rely on black-box treatment recommendation algorithms. The variable importance measures we consider are…
Sinan Acemoglu, Christian Kleiber, J. Urban
Variable importance in regression analyses is of considerable interest in a variety of fields. There is no unique method for assessing variable importance. However, a substantial share of the available literature employs Shapley values, either explicitly or implicitly, to decompose a suitable goodness-of-fit measure…
Guancheng Zhou, Haiping Xu, Jason Liu, Donghui Yan
Variable importance produced by Random Forests (RF) is used widely in statistical data analysis, and has played an important role in a variety of tasks such as assisting model interpretation, model selection and diagnosis, and cost-bounded learning etc. However, the calculation of variable importance in RF does not…
Yucheng Zhao, Brian D. Williamson
Variable importance may describe either intrinsic predictive information in a population or extrinsic importance for a fitted prediction rule. Quantifying the uncertainty in variable importance estimates is critical for interpretation. Methods for estimating intrinsic variable importance (we will refer to these as…
Zhikang Liu, Yiyang Niu, Tian Le, Daniel G Chen + 2 more
The rapid maturation of single-cell multi-omics technologies has enabled unprecedented resolution for mapping disease states and identifying disease-associated biomarkers. In practice, biomarkers are often discovered through differential detection that treat genomic features as independent contributors to phenotypes…
Tim Müller, Roman Hornung, Silke Szymczak, Hannes Buchner
Background Identifying relevant biomarkers is critical in clinical research and precision medicine, particularly when analysing high-dimensional data. Random forests (RFs) are promising for such settings due to their flexibility, ease of use, and their ability to handle data sets with more variables than samples. RFs…
Amr Mohamed, Kevin H. Lee
As data complexity and volume increase rapidly, efficient statistical methods for identifying significant variables become crucial. Variable selection plays a vital role in establishing relationships between predictors and response variables. The challenge lies in achieving this goal while controlling the False…
Authors not listed
Background: Batch reactor process optimization has traditionally relied on Analysis of Variance (ANOVA) for factor effect quantification. However, Structural Equation Modeling (SEM) and machine learning (ML) offer complementary mechanistic and predictive capabilities that remain underexplored in chemical engineering…
Bambang Widjanarko Otok, Zulfani Alfasanah, Diaz Fitra Aksioma
Title: Highlights 1. • PCA is used in the inner weighting scheme of the PLS model to obtain latent variable scores. 2. • PLS-IPA explains the influence between latent variables and indicators while mapping indicators based on importance and performance.
Jörn Lötsch, André Himmelspach, Dario Kringel
Title: Summary Generative AI can expand small biomedical datasets but may amplify noise and distort statistical relationships. We developed genESOM, a framework integrating an error control system into a generative AI method based on emergent self-organizing maps. By separating structure learning from data synthesis…
Anand Singh, Luke Pennella, Eshan Kabir, Xiaoxi Shen
Deep neural networks have been widely used in many applications (e.g., computer vision and natural language processing); however, understanding their explainability remains a challenging task. Recently, substantial research has been devoted to improving the explainability of deep neural networks, with most of this work…
Raelynn Chen, Attri Ghosh, Jie Hu, Yong Chen + 2 more
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does…
Keming Zhang, Yaoyao Li, Jungang Zou, Sijian Wang + 2 more
Selecting important individual- and cluster-level predictors has become increasingly critical in healthcare research, where data often exhibit hierarchical structures due to collection from multiple clusters. Mixed-effects models, which account for within-cluster correlation and between-cluster heterogeneity, are a…
Raelynn Chen, Attri Ghosh, Jie Hu, Yong Chen + 2 more
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does…
Ghislain Fievet, Julien Broséus, David Meyre, Sébastien Hergalant
In this study, we present a comprehensive evaluation framework for comparing various combinations of artificial intelligence (AI) methods in the context of explainable AI (XAI) for variable selection in experimental biological and biomedical data. Our goal was to assess the efficiency, computational cost, and accuracy…
Authors not listed
This research investigates predicting the Highest Occupied Molecular Orbital and the Lowest Unoccupied Molecular Orbital (HOMO-LUMO; short HL) gap of natural compounds, a crucial property for understanding molecular electronic behavior relevant to cheminformatics and materials science. To address the high computational…
Johannes A. Vey, Georg Heinze, Meinhard Kieser
Purpose A wide range of methods exist for developing a clinical prediction model (CPM) and for performing variable selection. Our purpose was to develop a fair simulation study design and to investigate the properties, strengths, and weaknesses of different methods to predict a continuous outcome in low-dimensional…
Authors not listed
Crystal structure prediction (CSP) is a valuable computational technique used to anticipate the likely crystal structures of a compound of interest. These methods have been proven useful in research and development of pharmaceutical solid forms and in guiding the discovery of materials with targeted properties. Despite…
Authors not listed
High-throughput experimentation (HTE) in materials science generates vast, high-dimensional datasets relating synthesis parameters to material properties. While machine learning (ML) models excel at predicting properties from these parameters, they often fail to distinguish causal drivers from merely correlated…
Erik D. VonKaenel, Lisa M. Bramer, Javier E. Flores, Thomas O Metz + 2 more
In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Authors not listed
The rapid growth of worldwide computing power has transformed in silico chemistry into a discipline that is integrated into the daily work of many chemists. Nowadays, researchers find it increasingly straightforward to predict a wide range of molecular properties and chemi- cal processes at reasonable computational…