21 papers · ranked by Valyu relevance
Scott Lundberg, Gabriel Erion, Su‐In Lee
Interpreting predictions from tree ensemble methods such as gradient boosting machines and random forests is important, yet feature attribution for trees is often heuristic and not individualized for each prediction. Here we show that popular feature attribution methods are inconsistent, meaning they can lower a…
Raquel Rodríguez-Pérez, Jürgen Bajorath
Difficulties in interpreting machine learning (ML) models and their predictions limit the practical applicability of and confidence in ML in pharmaceutical research. There is a need for agnostic approaches aiding in the interpretation of ML models regardless of their complexity that is also applicable to deep neural…
Rory Mitchell, Eibe Frank, Geoffrey Holmes, Alberto Cano
SHapley Additive exPlanation (SHAP) values ([24]) provide a game theoretic interpretation of the predictions of machine learning models based on Shapley values ([35]). While exact calculation of SHAP values is computationally intractable in general, a recursive polynomial-time algorithm called TreeShap ([23]) is…
Jilei Yang
SHAP (SHapley Additive exPlanation) values are one of the leading tools for interpreting machine learning models, with strong theoretical guarantees (consistency, local accuracy) and a wide availability of implementations and use cases. Even though computing SHAP values takes exponential time in general, TreeSHAP takes…
Olatomiwa O. Bifarin, Imran Ashraf
Machine learning (ML) models are used in clinical metabolomics studies most notably for biomarker discoveries, to identify metabolites that discriminate between a case and control group. To improve understanding of the underlying biomedical problem and to bolster confidence in these discoveries, model interpretability…
Thi-Thu-Huong Le, Haeyoung Kim, Hyoeun Kang, Howon Kim + 2 more
'Zhongyun Hua' 'Yushu Zhang'] In recent years, many methods for intrusion detection systems (IDS) have been designed and developed in the research community, which have achieved a perfect detection rate using IDS datasets. Deep neural networks (DNNs) are representative examples applied widely in IDS. However, DNN…
Tomohiro Ishibashi, Akio Onogi
Mapping quantitative trait loci (QTLs) is one of the major goals of quantitative genetics; however, identifying the interactions between QTLs remains challenging. Recently developed machine learning methods, such as deep learning and gradient boosting, are transforming the real world. These methods could advance QTL…
Peng Yu, Chao Xu, Albert Bifet, Jesse Read
Decision trees are well-known due to their ease of interpretability. To improve accuracy, we need to grow deep trees or ensembles of trees. These are hard to interpret, offsetting their original benefits. Shapley values have recently become a popular way to explain the predictions of tree-based machine learning models.…
Olatomiwa O. Bifarin
Machine learning (ML) models are used in clinical metabolomics studies most notably for biomarker discoveries, to identify metabolites that discriminate between a case and control group. To improve understanding of the underlying biomedical problem and to bolster confidence in these discoveries, model interpretability…
Pål V. Johnsen, Signe Riemer-Sørensen, Andrew Thomas DeWan, Megan E. Cahill + 1 more
'Megan E. Cahill' 'Mette Langaas'] Background The identification of gene-gene and gene-environment interactions in genome-wide association studies is challenging due to the unknown nature of the interactions and the overwhelmingly large number of possible combinations. Parametric regression models are suitable to look…
Pengyu Liu, Matthew Gould, Caroline Colijn
Phylogenetic trees are a central tool in evolutionary biology. They demonstrate evolutionary patterns among species, genes, and with modern sequencing technologies, patterns of ancestry among sets of individuals. Phylogenetic trees usually consist of tree shapes, branch lengths and partial labels. Comparing tree shapes…
Dominik Seidel, Peter Annighöfer, Melissa Stiers, Clara Delphine Zemp + 6 more
'Clara Delphine Zemp' 'Katharina Burkardt' 'Martin Ehbrecht' 'Katharina Willim' 'Holger Kreft' 'Dirk Hölscher' 'Christian Ammer'] Title: Abstract Aboveground tree architecture is neither fully deterministic nor random. It is likely the result of mechanisms that balance static requirements and light-capturing…
C. Colijn, G. Plazzotta
The shapes of evolutionary trees are influenced by the nature of the evolutionary process, but comparisons of trees from different processes are hindered by the challenge of completely describing tree shape. We present a full characterization of the shapes of rooted branching trees in a form that lends itself to…
Masrur Sobhan, Ananda Mohan Mondal
Lung cancer is the leading cause of cancer compared to other cancers in the USA despite being the most commonly diagnosed. The overall survival rate of lung cancer is not satisfactory even though having cutting edge treatment methods for cancers. Genomic profiling and biomarker gene identification of lung cancer…
Éric Hoffbeck, Ieke Moerdijk
We discuss a notion of shuffle for trees which extends the usual notion of a shuffle for two natural numbers. We give several equivalent descriptions, and prove some algebraic and combinatorial properties. In addition, we characterize shuffles in terms of open sets in a topological space associated to a pair of trees.…
Jirui Jin, Somayeh Faraji, Bin Liu, Mingjie Liu
Perovskite materials, renowned for their versatility and remarkable properties, pose challenges in discovering optimal candidates due to the vast compositional space. Data-driven machine learning (ML) offers promise in expediting material discovery; however, the trade-off between accuracy and efficiency across…
Authors not listed
Highly fluorinated aromatic compounds exhibit unique electronic structures, however their selective transformation remains a longstanding challenge. Halogenation of F7 naphthalene previously required low temperatures (–40 to 0 °C) for high yields, while room-temperature reactions suffered from side reactions and…
Authors not listed
We present a simple yet efficient random (brute-force) algorithm for constructing solvated molecular systems. By placing solvent molecules at random positions and orientations within a simulation box, we circumvent the complexities typically associated with more sophisticated packing algorithms. The main computational…
Jonas Schaub, Julian Zander, Achim Zielesny, Christoph Steinbeck
The concept of molecular scaffolds as defining core structures of organic molecules is utilised in many areas of chemistry and cheminformatics, e.g. drug design, chemical classification, or the analysis of high-throughput screening data. Here, we present Scaffold Generator, a comprehensive open library for the…
Sean Cleary, Mareike Fischer, Robert Griffiths, Raazesh Sainudiin
> Abstract. We introduce some natural families of distributions on rooted binary ranked plane trees with a view toward unifying ideas from various fields, including macroevolution, epidemiology, computational group theory, search algorithms and other fields. In the process we introduce the notions of…
Vladimir Kondratyev, Marian Dryzhakov, Timur Gimadiev, Dmitriy Slutskiy
In this work, we provide further development of the junction tree variational autoencoder (JT VAE) architecture in terms of implementation and application of the internal feature space of the model. Pretraining of JT VAE on a large dataset and further optimization with a regression model led to a latent space that can…