26 papers · ranked by Valyu relevance
Parisa Amin
At the time of diagnosis for cancer patients, a wide array of data can be gathered, ranging from clinical information to multiple layers of omics data. Determining which of these data are most informative is crucial, not only for advancing biological understanding but also for clinical and economic considerations. This…
Luke Power, Krishnendu Guha
—Many Machine Learning (ML) models are referred to as black-box models, providing no real insights into why a prediction is made. Feature importance and explainability are important for increasing transparency and trust in ML models, particularly in settings such as healthcare and finance. With quantum computing's…
Md Atik Bhuiyan, Md Rashik Shahriar Akash, Radiful Islam, Shohidul Islam Polash + 2 more
Dengue fever presents a growing public health challenge in tropical and subtropical regions, where early detection is crucial for effective intervention. This study conducts a comprehensive comparative analysis of 13 machine learning and deep learning models for nonclinical, symptom-based dengue prediction, focusing on…
Zuzanna Karwowska, Oliver Aasmets, Tomasz Kosciolek, Elin Org
Accurate classification of host phenotypes from microbiome data is essential for future therapies in microbiome-based medicine and machine learning approaches have proved to be an effective solution for the task. The complex nature of the gut microbiome, data sparsity, compositionality and population-specificity…
Hiromasa Kaneko
In molecular design, material design, process design, and process control, it is important not only to construct a model with high predictive ability between explanatory features x and objective features y using a dataset but also to interpret the constructed model. An index of feature importance in x is permutation…
Shadi Jacob Khoury, Yazeed Zoabi, Mickey Scheinowitz, Noam Shomron + 1 more
'Ronald N. Harty'] In this study, we introduce a novel approach that integrates interpretability techniques from both traditional machine learning (ML) and deep neural networks (DNN) to quantify feature importance using global and local interpretation methods. Our method bridges the gap between interpretable ML models…
Temesgen Mehari, Ashish Sundar, Alen Bošnjaković, P. Harris + 6 more
'Steven E. Williams' 'Axel Loewe' 'Olaf Doessel' 'Claudia Nagel' 'Nils Strodthoff' 'Philip J. Aston'] Abstract—Feature importance methods promise to provide a ranking of features according to importance for a given classification task. A wide range of methods exist but their rankings often disagree and they are…
Terence Parr, James Wilson, Jeff Hamrick
Practitioners use feature importance to rank and eliminate weak predictors during model development in an effort to simplify models and improve generality. Unfortunately, they also routinely conflate such feature importance measures with feature impact, the isolated effect of an explanatory variable on the response…
Pål Vegard Johnsen, Inga Strümke, Mette Langaas, Andrew Thomas DeWan + 2 more
Estimating feature importance, which is the contribution of a prediction or several predictions due to a feature, is an essential aspect of explaining data-based models. Besides explaining the model itself, an equally relevant question is which features are important in the underlying data generating process. We…
Eloisa Rocha Liedl, Shabeer Mohamed Yassin, Melpomeni Kasapi, Joram M. Posma
Cancer is the second leading cause of disease-related death worldwide, and machine learning-based identification of novel biomarkers is crucial for improving early detection and treatment of various cancers. A key challenge in applying machine learning to high-dimensional data is deriving important features in an…
Divish Rengasamy, Jimiama Mafeni Mase, Mercedes Torres Torres, Benjamin Rothwell + 2 more
'Benjamin Rothwell' 'David A. Winkler' 'Grazziela P. Figueredo'] Abstract—With the widespread use of machine learning to support decision-making, it is increasingly important to verify and understand the reasons why a particular output is produced. Although post-training feature importance approaches assist this…
Stephane Doyen, Hugh Taylor, Peter Nicholas, Lewis Crawford + 3 more
As an alternative, partial dependence plots (PDP) are often used to visualize decision boundaries. They help to describe relationships with non-linear effects, and show interactions between features. To construct such a plot, the PDP function is calculated at each possible value of a feature, representing the average…
Suvo Banik, Karthik Balasubramanian, Sukriti Manna, Sybil Derrible + 1 more
Identifying key descriptors and understanding important features across different classes of materials are crucial for machine learning (ML) tools to both predict material properties and reveal the physics underlying any process of interest. Traditionally, the predictive modeling of elastic properties of materials is…
Amnon Catav, Boyang Fu, Jason B. Ernst, Sriram Sankararaman + 1 more
The Natural Case Authors: ['Amnon Catav' 'Boyang Fu' 'Jason B. Ernst' 'Sriram Sankararaman' 'Ran Gilad-Bachrach'] When training a predictive model over medical data, the goal is sometimes to gain insights about a certain disease. In such cases, it is common to use feature importance as a tool to highlight significant…
Charles Westphal, Stephen Hailes, Mirco Musolesi
Selection Authors: ['Charles Westphal' 'Stephen Hailes' 'Mirco Musolesi'] In this paper, we introduce Partial Information Decomposition of Features (PIDF), a new paradigm for simultaneous data interpretability and feature selection. Contrary to traditional methods that assign a single importance value, our approach is…
Authors not listed
Highly fluorinated aromatic compounds exhibit unique electronic structures, however their selective transformation remains a longstanding challenge. Halogenation of F7 naphthalene previously required low temperatures (–40 to 0 °C) for high yields, while room-temperature reactions suffered from side reactions and…
Jianqiang Sheng, Songhua Xu, Xiaonan Luo
representation Authors: ['Jianqiang Sheng' 'Songhua Xu' 'Xiaonan Luo'] Background Images embedded in biomedical publications carry rich information that often concisely summarize key hypotheses adopted, methods employed, or results obtained in a published study. Therefore, they offer valuable clues for understanding…
Leif Hancox-Li, I. Elizabeth Kumar
As the public seeks greater accountability and transparency from machine learning algorithms, the research literature on methods to explain algorithms and their outputs has rapidly expanded. Feature importance methods form a popular class of explanation methods. In this paper, we apply the lens of feminist epistemology…
Jianzhong Chen, Leon Qi Rong Ooi, Trevor Wei Kiat Tan, Shaoshi Zhang + 6 more
There is significant interest in using neuroimaging data to predict behavior. The predictive models are often interpreted by the computation of feature importance, which quantifies the predictive relevance of an imaging feature. 62 suggest that feature importance estimates exhibit low split-half reliability, as well as…
Ye Tian, Andrew Zalesky
Cognitive performance can be predicted from an individual’s functional brain connectivity with modest accuracy using machine learning approaches. As yet, however, predictive models have arguably yielded limited insight into the neurobiological processes supporting cognition. To do so, feature selection and feature…
Authors not listed
Quantification is a challenge for non-targeted analysis (NTA) with liquid chromatography–high resolution mass spectrometry (LC–HRMS), due to the lack of analytical standards. Quantification via structure-based predicted ionization efficiency (IE) was found to provide the highest accuracy in estimating concentration.…
Authors not listed
Traditional and non-classical machine learning models for solid-state structure prediction have predominantly relied on compositional features (derived from properties of constituent elements) to predict the existence of structure and its properties. However, the lack of structural information can be a source of…
Sang-Il Choi, Su-Hyun Kim, Yoonseok Yang, Gu-Min Jeong
We propose a data refinement and channel selection method for vapor classification in a portable e-nose system. For the robust e-nose system in a real environment, we propose to reduce the noise in the data measured by sensor arrays and distinguish the important part in the data by the use of feature feedback.…
Simon Viet Johansson, Hampus Gummesson Svensson, Esben Bjerrum, Alexander Schliep + 3 more
Computer aided synthesis planning is a rapidly growing field for suggesting synthetic routes for molecules of interest. The methods used are usually dependent on access to large datasets for training, but with a finite experimental budget there are limitations on how much data can be obtained from experiments. Active…
Devi Ganapathi, Wunmi Akinlemibola, Antonio Baclig, Emily Penn + 1 more
Quinones and hydroquinones are small organic molecules with numerous applications: battery electrolytes, pharmaceuticals, sensors, to name a few. An understanding of their fundamental properties, such as melting points, is essential to incorporate these compounds into relevant technologies. In this study, two different…
Authors not listed
Learning aqueous solubility remains a key challenge in drug development for improving oral bioavailability. Traditional data-driven solubility estimations using standard supervised models, however, can often suppress the information embedded in a molecule’s chemical properties and the intricate connectivity of its…