19 papers · ranked by Valyu relevance
Zoran Bursac, C Heath Gauss, David Keith Williams, David W Hosmer
Background The main problem in many model-building situations is to choose from a large set of covariates those that should be included in the "best" model. A decision to keep a variable in the model might be based on the clinical or statistical significance. There are several variable selection algorithms in…
Elizabeth Handorf, Yinuo Yin, Michael Slifker, Shannon Lynch
Background Social-environmental data obtained from the US Census is an important resource for understanding health disparities, but rarely is the full dataset utilized for analysis. A barrier to incorporating the full data is a lack of solid recommendations for variable selection, with researchers often hand-selecting…
Eliana Lima, Peers Davies, Jasmeet Kaler, Fiona Lovatt + 1 more
Variable selection in inferential modelling is problematic when the number of variables is large relative to the number of data points, especially when multicollinearity is present. A variety of techniques have been described to identify ‘important’ subsets of variables from within a large parameter space but these may…
Donald Douglas Atsa'am, Ruth Wario, Pakiso Khomokhoana
1## Introduction In modelling, the concept of variable selection refers to the process of choosing a subset from the domain of input variables that result in models with good fit and high accuracy. It has been pointed out that not all the variables that make up the data set of a problem domain convey information of…
A. D. V. Tharkeshi T. Dharmaratne, Alysha De Livera, Stelios Georgiou, Stella Stylianou + 1 more
'Stella Stylianou' 'Mahdi Roozbeh'] Variable selection methods are widely used in observational studies. While many penalty-based statistical methods introduced in recent decades have primarily focused on prediction, classical statistical methods remain the standard approach in applied research and education. In this…
Willi Sauerbrei, Aris Perperoglou, Matthias Schmid, Michal Abrahamowicz + 6 more
'Michal Abrahamowicz' 'Heiko Becher' 'Harald Binder' 'Daniela Dunkler' 'Frank E. Harrell Jr' 'Patrick Royston' 'Georg Heinze' ''] Background How to select variables and identify functional forms for continuous variables is a key concern when creating a multivariable model. Ad hoc ‘traditional’ approaches to variable…
Qing-Yan Yin, Jun-Li Li, Chun-Xia Zhang
As a pivotal tool to build interpretive models, variable selection plays an increasingly important role in high-dimensional data analysis. In recent years, variable selection ensembles (VSEs) have gained much interest due to their many advantages. Stability selection (Meinshausen and Bühlmann, 2010), a VSE technique…
Theresa Ullmann, Georg Heinze, Lorena Hafermann, Christine Schilhart-Wallisch + 2 more
'Christine Schilhart-Wallisch' 'Daniela Dunkler' '' 'Suyan Tian'] Researchers often perform data-driven variable selection when modeling the associations between an outcome and multiple independent variables in regression analysis. Variable selection may improve the interpretability, parsimony and/or predictive…
Daniela Dunkler, Max Plischke, Karen Leffondré, Georg Heinze + 1 more
'Jake Olivier'] Statistical models are simple mathematical rules derived from empirical data describing the association between an outcome and several explanatory variables. In a typical modeling situation statistical analysis often involves a large number of potential explanatory variables and frequently only partial…
Pi Guo, Fangfang Zeng, Xiaomin Hu, Dingmei Zhang + 4 more
The variable selection technique is employed for epidemiologic analysis to identify independent associations between collective exposures and a health outcome . Selection of the best variables is aimed at controlling confounders to obtain unbiased estimates of covariate effects and predicting probabilities with robust…
Jacob Seedorff, Joseph E. Cavanaugh, Ciprian Giurcaneanu
One of the primary issues that arises in statistical modeling pertains to the assessment of the relative importance of each variable in the model. A variety of techniques have been proposed to quantify variable importance for regression models. However, in the context of best subset selection, fewer satisfactory…
Bethany J. Wolf, Yunyun Jiang, Sylvia H. Wilson, Jim C. Oates
There are several methods proposed in the literature for identifying the best subset of predictors in a GLMM or GEE setting for a repeatedly measured binary response. In this section, we describe these methods in greater detail and discuss the advantages and disadvantages of each method. Stepwise algorithms iteratively…
Christina Brester, Jussi Kauhanen, Tomi-Pekka Tuomainen, Sari Voutilainen + 4 more
'Sari Voutilainen' 'Mauno Rönkkö' 'Kimmo Ronkainen' 'Eugene Semenkin' 'Mikko Kolehmainen'] Background The redundancy of information is becoming a critical issue for epidemiologists. High-dimensional datasets require new effective variable selection methods to be developed. This study implements an advanced evolutionary…
Nathaniel S O’Connell, Byron C Jaeger, Garrett S Bullock, Jaime Lynn Speiser
'Jaime Lynn Speiser'] Title: Abstract Random forest (RF) regression is popular machine learning method to develop prediction models for continuous outcomes. Variable selection, also known as feature selection or reduction, involves selecting a subset of predictor variables for modeling. Potential benefits of variable…
Dake Hou, Wenli Zhou, Qiuxia Zhang, Kun Zhang + 2 more
'Xiangtao Li'] This study employs the principles of computer science and statistics to evaluate the efficacy of the linear random effect model, utilizing Lasso variable selection techniques (including Lasso, Elastic-Net, Adaptive-Lasso, and SCAD) through numerical simulation and empirical research. The analysis focuses…
Maryam Sadiq, Nasser A. Alsadhan, Ramla Shah, Sidra Younas + 2 more
'Zahid Rasheed' 'Suyan Tian'] Variable selection methods are very popular, especially in the field of big data with large predictors. These procedures improve the accuracy and performance of the model by eliminating irrelevant and redundant variables. The main contribution of this study is to couple a logit model with…
C. Fernandez-Lozano, C. Canto, M. Gestal, J. M. Andrade-Garda + 3 more
'J. R. Rabuñal' 'J. Dorado' 'A. Pazos'] Given the background of the use of Neural Networks in problems of apple juice classification, this paper aim at implementing a newly developed method in the field of machine learning: the Support Vector Machines (SVM). Therefore, a hybrid model that combines genetic algorithms…
Shuzhen Sun, Zhuqi Miao, Blaise Ratcliffe, Polly Campbell + 5 more
With the rapid advancement of DNA sequencing technology, the volume and dimension of biological and medical data have been increasing at an unprecedented rate. Accompanying such high volume genetic data, the ‘curse of dimensionality’ has challenged the validity of statistical methods that do not scale to massive data.…
Eliana Lima, Robert Hyde, Martin Green
Inferential research commonly involves identification of causal factors from within high dimensional data but selection of the ‘correct’ variables can be problematic. One specific problem is that results vary depending on statistical method employed and it has been argued that triangulation of multiple methods is…