28 papers · ranked by Valyu relevance
Zoran Bursac, C Heath Gauss, David Keith Williams, David W Hosmer
Background The main problem in many model-building situations is to choose from a large set of covariates those that should be included in the "best" model. A decision to keep a variable in the model might be based on the clinical or statistical significance. There are several variable selection algorithms in…
Eliana Lima, Peers Davies, Jasmeet Kaler, Fiona Lovatt + 1 more
Variable selection in inferential modelling is problematic when the number of variables is large relative to the number of data points, especially when multicollinearity is present. A variety of techniques have been described to identify ‘important’ subsets of variables from within a large parameter space but these may…
Marika Ström, Nicole Wagner, Iryna Kolosenko, Åsa M. Wheelock
The R-workflow ropls-ViPerSNet (R orthogonal projections of latent structures with Variable Permutation Selection and Elastic Net) facilitates variable selection, model optimization and significance testing using permutations of OPLS-DA models, with the scaled loadings (p[corr]) as the main metric of significance…
Roberta Coletti, J. Orestes Cerdeira, Marcos Raydan, Marta B. Lopes
High-dimensional omics data often contain more variables than observations, which negatively impacts the performance of classical data analysis methods. Dimensionality reduction is typically addressed through variable selection strategies that incorporate a penalty term into the model. While effective for selecting…
Adriano Zanin Zambom, Michael G. Akritas
Let X, Z be r and s-dimensional covariates, respectively, used to model the response variable Y as Y = m(X,Z) + σ(X,Z)ǫ. We develop an ANOVA-type test for the null hypothesis that Z has no influence on the regression function, based on residuals obtained from local polynomial fitting of the null model. Using p-values…
A. D. V. Tharkeshi T. Dharmaratne, Alysha De Livera, Stelios Georgiou, Stella Stylianou + 1 more
'Stella Stylianou' 'Mahdi Roozbeh'] Variable selection methods are widely used in observational studies. While many penalty-based statistical methods introduced in recent decades have primarily focused on prediction, classical statistical methods remain the standard approach in applied research and education. In this…
Chuji Luo, Michael J. Daniels
Variable selection is an important statistical problem. This problem becomes more challenging when the candidate predictors are of mixed type (e.g. continuous and binary) and impact the response variable in nonlinear and/or non-additive ways. In this paper, we review existing variable selection approaches for the…
Gao Wang, Abhishek Sarkar, Peter Carbonetto, Matthew Stephens
We introduce a simple new approach to variable selection in linear regression, and to quantifying uncertainty in selected variables. The approach is based on a new model – the “Sum of Single Effects” (SuSiE) model – which comes from writing the sparse vector of regression coefficients as a sum of “single-effect”…
Willi Sauerbrei, Aris Perperoglou, Matthias Schmid, Michal Abrahamowicz + 6 more
'Michal Abrahamowicz' 'Heiko Becher' 'Harald Binder' 'Daniela Dunkler' 'Frank E. Harrell Jr' 'Patrick Royston' 'Georg Heinze' ''] Background How to select variables and identify functional forms for continuous variables is a key concern when creating a multivariable model. Ad hoc ‘traditional’ approaches to variable…
Theresa Ullmann, Georg Heinze, Lorena Hafermann, Christine Schilhart-Wallisch + 2 more
'Christine Schilhart-Wallisch' 'Daniela Dunkler' '' 'Suyan Tian'] Researchers often perform data-driven variable selection when modeling the associations between an outcome and multiple independent variables in regression analysis. Variable selection may improve the interpretability, parsimony and/or predictive…
Daniela Dunkler, Max Plischke, Karen Leffondré, Georg Heinze + 1 more
'Jake Olivier'] Statistical models are simple mathematical rules derived from empirical data describing the association between an outcome and several explanatory variables. In a typical modeling situation statistical analysis often involves a large number of potential explanatory variables and frequently only partial…
Pi Guo, Fangfang Zeng, Xiaomin Hu, Dingmei Zhang + 4 more
The variable selection technique is employed for epidemiologic analysis to identify independent associations between collective exposures and a health outcome . Selection of the best variables is aimed at controlling confounders to obtain unbiased estimates of covariate effects and predicting probabilities with robust…
Jeffrey L. Andrews, Paul D. McNicholas
As data sets continue to grow in size and complexity, effective and efficient techniques are needed to target important features in the variable space. Many of the variable selection techniques that are commonly used alongside clustering algorithms are based upon determining the best variable subspace according to…
Maryam Sadiq, Nasser A. Alsadhan, Ramla Shah, Sidra Younas + 2 more
'Zahid Rasheed' 'Suyan Tian'] Variable selection methods are very popular, especially in the field of big data with large predictors. These procedures improve the accuracy and performance of the model by eliminating irrelevant and redundant variables. The main contribution of this study is to couple a logit model with…
Insha Ullah, Kerrie Mengersen, Anthony Pettitt, Benoit Liquet
High-dimensional datasets, where the number of variables ‘p’ is much larger compared to the number of samples ‘n’, are ubiquitous and often render standard classification and regression techniques unreliable due to overfitting. An important research problem is feature selection — ranking of candidate variables based on…
Nilotpal Sanyal
High-dimensional data with binary outcomes are common in biological sciences and related fields such as omics, epidemiology, and healthcare, environmental sciences, physical and engineering sciences, and business and policy making. Often, such data are sparse so that only a small number of all available variables truly…
Zhou Tang, Ted Westling
Selecting from or ranking a set of candidates variables in terms of their capacity for predicting an outcome of interest is an important task in many scientific fields. A variety of methods for variable selection and ranking have been proposed in the literature. In practice, it can be challenging to know which method…
Shuzhen Sun, Zhuqi Miao, Blaise Ratcliffe, Polly Campbell + 4 more
High-throughput sequencing technology has revolutionized both medical and biological research by generating exceedingly large numbers of genetic variants. The resulting datasets share a number of common characteristics that might lead to poor generalization capacity. Concerns include noise accumulated due to the large…
Souvik Bag, Kapil Gupta, Soudeep Deb
The selection of essential variables in logistic regression is vital because of its extensive use in medical studies, finance, economics and related fields. In this paper, we explore four main typologies (test-based, penalty-based, screening-based, and tree-based) of frequentist variable selection methods in logistic…
Raelynn Chen, Attri Ghosh, Jie Hu, Yong Chen + 2 more
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does…
Raelynn Chen, Attri Ghosh, Jie Hu, Yong Chen + 2 more
High-dimensional biomedical datasets routinely contain sparse signals embedded among vast, correlated features, making variable selection central to building models that generalize. Although significance-based selection is widely used across modalities (e.g., imaging, EHR, multi-omics), statistical significance does…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Rui Liu, Nan Sun
Selecting important features in high-dimensional survival analysis is critical for identifying confirmatory biomarkers while maintaining rigorous error control. In this paper, we propose a derandomized knockoffs procedure for Cox regression that enhances stability in feature selection while maintaining rigorous control…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Authors not listed
High-throughput experimentation (HTE) in materials science generates vast, high-dimensional datasets relating synthesis parameters to material properties. While machine learning (ML) models excel at predicting properties from these parameters, they often fail to distinguish causal drivers from merely correlated…
James J. Cai
The recent development of single-cell technologies, especially single-cell RNA sequencing (scRNA-seq), provides an unprecedented level of resolution to the cell type heterogeneity. It also enables the study of gene expression variability across individual cells within a homogenous cell population. Feature selection…
Authors not listed
Experimental design plays an important role in efficiently acquiring informative data for system characterization and deriving robust conclusions under resource limitations. Recent advancements in high-throughput experimentation coupled with machine learning have notably improved experimental procedures. While Bayesian…
Authors not listed
Plastic mechanical recycling is the conventional technological step towards circularity. In such aspects, complex mixtures of polyolefin blends are often fed into mechanical recycling systems, resulting in moulded products with uncertain quality. To add to the difficulty of heterogeneous feedstocks, the testing of…