15 papers · ranked by Valyu relevance
Donald Douglas Atsa'am, Ruth Wario, Pakiso Khomokhoana
1## Introduction In modelling, the concept of variable selection refers to the process of choosing a subset from the domain of input variables that result in models with good fit and high accuracy. It has been pointed out that not all the variables that make up the data set of a problem domain convey information of…
A. D. V. Tharkeshi T. Dharmaratne, Alysha De Livera, Stelios Georgiou, Stella Stylianou + 1 more
'Stella Stylianou' 'Mahdi Roozbeh'] Variable selection methods are widely used in observational studies. While many penalty-based statistical methods introduced in recent decades have primarily focused on prediction, classical statistical methods remain the standard approach in applied research and education. In this…
Theresa Ullmann, Georg Heinze, Lorena Hafermann, Christine Schilhart-Wallisch + 2 more
'Christine Schilhart-Wallisch' 'Daniela Dunkler' '' 'Suyan Tian'] Researchers often perform data-driven variable selection when modeling the associations between an outcome and multiple independent variables in regression analysis. Variable selection may improve the interpretability, parsimony and/or predictive…
Dongsheng Li, Chunyan Pan, Jing Zhao, Anfei Luo + 1 more
This paper presents a multi-algorithm fusion model (StackingGroup) based on the Stacking ensemble learning framework to address the variable selection problem in high-dimensional group structure data. The proposed algorithm takes into account the differences in data observation and training principles of different…
Yuan Zhou, Botao Fa, Ting Wei, Jianle Sun + 2 more
'Yue Zhang'] Investigation of the genetic basis of traits or clinical outcomes heavily relies on identifying relevant variables in molecular data. However, characteristics such as high dimensionality and complex correlation structures of these data hinder the development of related methods, resulting in the inclusion…
Jacob Seedorff, Joseph E. Cavanaugh, Ciprian Giurcaneanu
One of the primary issues that arises in statistical modeling pertains to the assessment of the relative importance of each variable in the model. A variety of techniques have been proposed to quantify variable importance for regression models. However, in the context of best subset selection, fewer satisfactory…
Yonghan Kwon, Kyunghwa Han, Young Joo Suh, Inkyung Jung
Stability selection is a variable selection algorithm based on resampling a dataset. Based on stability selection, we propose weighted stability selection to select variables by weighing them using the area under the receiver operating characteristic curve (AUC) from additional modelling. Through an extensive…
Dake Hou, Wenli Zhou, Qiuxia Zhang, Kun Zhang + 2 more
'Xiangtao Li'] This study employs the principles of computer science and statistics to evaluate the efficacy of the linear random effect model, utilizing Lasso variable selection techniques (including Lasso, Elastic-Net, Adaptive-Lasso, and SCAD) through numerical simulation and empirical research. The analysis focuses…
Nathaniel S O’Connell, Byron C Jaeger, Garrett S Bullock, Jaime Lynn Speiser
'Jaime Lynn Speiser'] Title: Abstract Random forest (RF) regression is popular machine learning method to develop prediction models for continuous outcomes. Variable selection, also known as feature selection or reduction, involves selecting a subset of predictor variables for modeling. Potential benefits of variable…
Maryam Sadiq, Nasser A. Alsadhan, Ramla Shah, Sidra Younas + 2 more
'Zahid Rasheed' 'Suyan Tian'] Variable selection methods are very popular, especially in the field of big data with large predictors. These procedures improve the accuracy and performance of the model by eliminating irrelevant and redundant variables. The main contribution of this study is to couple a logit model with…
Baidu Li, Xinhai Li
Linear models, including t-test, ANOVA, regression, ANCOVA, and generalized linear models, are foundational tools in statistical analysis. For large datasets, such as those involving tens of thousands of genes and millions of records, numerous advanced methods have been developed to improve both computational…
Linard Hoessly, Jaromil Frossard, Simon Schwab, Frédérique Chammartin + 5 more
'Frédérique Chammartin' 'Alexander Leichtle' 'Peter Werner Schreiber' 'Dionysios Neofytos' 'Michael Koller' '' 'Syed Nisar Hussain Bukhari'] The integration of machine learning methodologies has become prevalent in the development of clinical prediction models, often suggesting superior performance compared to…
Edwin Kipruto, Willi Sauerbrei, Ralf Bender
In low-dimensional data and within the framework of a classical linear regression model, we intend to compare variable selection methods and investigate the role of shrinkage of regression estimates in a simulation study. Our primary aim is to build descriptive models that capture the data structure parsimoniously…
Michael Kammer, Daniela Dunkler, Stefan Michiels, Georg Heinze
Background Variable selection for regression models plays a key role in the analysis of biomedical data. However, inference after selection is not covered by classical statistical frequentist theory, which assumes a fixed set of covariates in the model. This leads to over-optimistic selection and replicability issues.…
Tim Müller, Roman Hornung, Silke Szymczak, Hannes Buchner
Background Identifying relevant biomarkers is critical in clinical research and precision medicine, particularly when analysing high-dimensional data. Random forests (RFs) are promising for such settings due to their flexibility, ease of use, and their ability to handle data sets with more variables than samples. RFs…