22 papers · ranked by Valyu relevance
Kathryn L Lunetta, L Brooke Hayward, Jonathan Segal, Paul Van Eerdewegh
'Paul Van Eerdewegh'] Background Genome-wide association studies for complex diseases will produce genotypes on hundreds of thousands of single nucleotide polymorphisms (SNPs). A logical first approach to dealing with massive numbers of SNPs is to use some test to screen the SNPs, retaining only those that meet some…
Charles D. Coleman
aCharles D. Coleman is a Mathematical Statistician, U.S. Census Bureau, CENHQ 5H48C, Washington, DC, 20233 (e-mail: charles.d.coleman@census.gov). Any views expressed are those of the author and not necessarily those of the U.S. Census Bureau. The Census Bureau has reviewed this data product for unauthorized disclosure…
Jacob Seedorff, Joseph E. Cavanaugh, Ciprian Giurcaneanu
One of the primary issues that arises in statistical modeling pertains to the assessment of the relative importance of each variable in the model. A variety of techniques have been proposed to quantify variable importance for regression models. However, in the context of best subset selection, fewer satisfactory…
Jeremy VanderDoes, Claire Marceaux, Kenta Yokote, Marie-Liesse Asselin-Labat + 3 more
'Marie-Liesse Asselin-Labat' 'Gregory Rice' 'Jack D. Hywood' 'Pedro Mendes'] Tumor microenvironments (TMEs) contain vast amounts of information on patient’s cancer through their cellular composition and the spatial distribution of tumor cells and immune cell populations. Exploring variations in TMEs between patient…
Yucheng Zhao, Brian D. Williamson
Variable importance may describe either intrinsic predictive information in a population or extrinsic importance for a fitted prediction rule. Quantifying the uncertainty in variable importance estimates is critical for interpretation. Methods for estimating intrinsic variable importance (we will refer to these as…
Chenglong Ye, Yi Yang, Yuhong Yang
With now well-recognized non-negligible model selection uncertainty, data analysts should no longer be satisfied with the output of a single final model from a model selection process, regardless of its sophistication. To improve reliability and reproducibility in model choice, one constructive approach is to make good…
Vanessa Gómez-Verdejo, Emilio Parrado-Hernández, Jussi Tohka
An important problem that hinders the use of supervised classification algorithms for brain imaging is that the number of variables per single subject far exceeds the number of training subjects available. Deriving multivariate measures of variable importance becomes a challenge in such scenarios. This paper proposes a…
Ignasi Arranz, Ralph Mac Nally, Emili García‐Berthou
Identifying the most important variables that determine patterns and processes is one of the main goals in many scientific fields, including ecological and evolutionary studies. Variable or relative importance is generally seen as the proportion of the variation in a response variable explained directly and indirectly…
Yilin Ning, Siqi Li, Yih Yng Ng, Michael Yih Chong Chia + 9 more
'Han Nee Gan' 'Ling Tiah' 'Desmond Renhao Mao' 'Wei Ming Ng' 'Benjamin Sieu-Hon Leong' 'Nausheen Doctor' 'Marcus Eng Hock Ong' 'Nan Liu' 'Po-Chih Kuo'] Machine learning (ML) methods are increasingly used to assess variable importance, but such black box models lack stability when limited in sample sizes, and do not…
Mohammad Kaviul Anam Khan, Rafal Kustra
In this paper we define a population parameter, "Generalized Variable Importance Metric (GVIM)", to measure importance of predictors for black box machine learning methods, where the importance is not represented by model-based parameter. GVIM is defined for each input variable, using the true conditional expectation…
Louis Mozart Kamdem, Ernest Fokoué
Estimating the importance of variables is an essential task in modern machine learning. This help to evaluate the goodness of a feature in a given model. Several techniques for estimating the importance of variables have been developed during the last decade. In this paper, we proposed a computational and theoretical…
Masayoshi Mase, Art B. Owen, Benjamin B. Seiler
The most popular methods for measuring importance of the variables in a black-box prediction algorithm make use of synthetic inputs that combine predictor variables from multiple observations. These inputs can be unlikely, physically impossible, or even logically impossible. As a result, the predictions for such cases…
Anne-Laure Boulesteix, Silke Janitza, Alexander Hapfelmeier, Kristel Van Steen + 1 more
'Kristel Van Steen' 'Carolin Strobl'] In an interesting and quite exhaustive review on Random Forests (RF) methodology in bioinformatics Touw et al. address-among other topics-the problem of the detection of interactions between variables based on RF methodology. We feel that some important statistical concepts, such…
Cesaré Ovando-Vázquez, Daniel Cázarez-García, Robert Winkler
Machine learning algorithms excavate important variables from biological big data. However, deciding on the biological relevance of identified variables is challenging. The addition of artificial noise, ‘decoy’ variables, to raw data, ‘target’ variables, enables calculating a false-positive rate (FPR) and a biological…
Robert Dunne
Random Forests (RF) are a very widely used modelling tool. 34 concludes that no nonlinear model had a more widespread popularity, from health care to academia to industry, than random forests and decision trees. The bounds of the methodology are still being extended. 4 give an example with 80 million variables. It is…
Adam B. Smith, Maria J. Santos
Models of species’ distributions and niches are frequently used to infer the importance of range- and niche-defining variables. However, the degree to which these models can reliably identify important variables and quantify their influence remains unknown. Here we use a series of simulations to explore how well models…
Authors not listed
Identifying the crystallographic positions of dopant ions in doped inorganic materials is a longstanding challenge that has impeded precise control over material properties. To address this, we introduce a robust theoretical framework along with detailed methodologies to evaluate dopant ion occupancy within a host…
Authors not listed
High-throughput experimentation (HTE) in materials science generates vast, high-dimensional datasets relating synthesis parameters to material properties. While machine learning (ML) models excel at predicting properties from these parameters, they often fail to distinguish causal drivers from merely correlated…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Anastasia Sholokhova, Dmitriy Matyushin, Mikhail Shashkov
Ionic liquids, i.e., organic salts with a low melting point, can be used as gas chromatographic liquid stationary phases. These stationary phases have some advantages such as peculiar selectivity, high polarity, and thermostability. Many previous works are devoted to such stationary phases. However, there are still no…
Koichi Handa, Sakae Sugiyama, Michiharu Kageyama, Takeshi Iijima
It is important to precisely predict the intestinal absorption ratio (Fa) at an early stage in the discovery of orally available drugs because it directly influences drug efficacy. Gastrointestinal unified theoretical framework (GUTFW) and machine learning (ML) are commonly used to predict the percentage of Fa. In…
Authors not listed
Meteorological normalization is a key concept in studying anthropogenic effects on air pollutant concentrations and its temporal trends. While apparently successful in revealing anthropogenic effects and often used, there are downsides to the methods and limitations which should be taken into account when using it.…