28 papers · ranked by Valyu relevance
Yanchao Liu
This paper proposes a new mixed-integer programming (MIP) formulation to optimize split rule selection in the decision tree induction process, and develops an efficient search algorithm that is able to solve practical instances of the MIP model faster than commercial solvers. The formulation is novel for it directly…
Víctor Francisco Sampedro Blanco, Alberto Japón, Justo Puerto
In this paper we present a novel mathematical optimizationbased methodology to construct tree-shaped classification rules for multiclass instances. Our approach consists of building Classification Trees in which, except for the leaf nodes, the labels are temporarily left out and grouped into two classes by means of a…
Wojciech Wieczorek, Jan Kozak, Łukasz Strąk, Arkadiusz Nowakowski + 1 more
A new two-stage method for the construction of a decision tree is developed. The first stage is based on the definition of a minimum query set, which is the smallest set of attribute-value pairs for which any two objects can be distinguished. To obtain this set, an appropriate linear programming model is proposed. The…
Helen L. Smith, Patrick J. Biggs, Nigel P. French, Adam N. H. Smith + 1 more
To date, there remains no satisfactory solution for absent levels in random forest models. Absent levels are levels of a predictor variable encountered during prediction for which no explicit rule exists. Imposing an order on nominal predictors allows absent levels to be integrated and used for prediction. The ordering…
Krzysztof Gajowniczek, Marcin Dudziński, Raúl Alcaraz, Luca Faes + 2 more
'Leandro Pardo' 'Boris Ryabko'] The primary objective of our study is to analyze how the nature of explanatory variables influences the values and behavior of impurity measures, including the Shannon, Rényi, Tsallis, Sharma-Mittal, Sharma-Taneja, and Kapur entropies. Our analysis aims to use these measures in the…
Simultaneous Latent Budget Trees for Stratified Classification Cristian Buoncompagni, Stefano Pellegrino, Giulia Vannucci, Roberta Siciliano
In the era of Explainable Artificial Intelligence, there is a renewed focus on single trees for their ease of interpretation. This paper introduces Simultaneous Latent Budget Trees, a probabilistic machine learning framework for classification trees in the presence of a stratification factor such as a temporal…
Samad Moslehi, Niloofar Rabiei, Ali Reza Soltanian, Mojgan Mamani
Background Due to the high mortality of COVID-19 patients, the use of a high-precision classification model of patient’s mortality that is also interpretable, could help reduce mortality and take appropriate action urgently. In this study, the random forest method was used to select the effective features in COVID-19…
Bin Yang, Wenzheng Bao, Jinglong Wang
Hypertension is a chronic disease and major risk factor for cardiovascular and cerebrovascular diseases that often leads to damage to target organs. The prevention and treatment of hypertension is crucially important for human health. In this paper, a novel ensemble method based on a flexible neural tree (FNT) is…
Y. Coadou
Boosted decision trees are a very powerful machine learning technique. After introducing specific concepts of machine learning in the highenergy physics context and describing ways to quantify the performance and training quality of classifiers, decision trees are described. Some of their shortcomings are then…
Nikola Anđelić, Sandi Baressi Šegota, Claudio Luparello
Simple Summary Breast cancer is a type of cancer with several sub-types and correct sub-type classification based on a large number of gene expressions is challenging even for artificial intelligence. However, the accurate classification of breast cancer in a patient is mandatory for the application of proper…
Min Lu, Ruijie Yin, X. Steven Chen
Building Single Sample Predictors (SSPs) from gene expression profiles presents challenges, notably due to the lack of calibration across diverse gene expression measurement technologies. However, recent research indicates the viability of classifying phenotypes based on the order of expression of multiple genes.…
Bhekisipho Twala, Eamon Molloy
An ensemble of classifiers combines several single classifiers to deliver a final prediction or classification decision. An increasingly provoking question is whether such an ensemble can outperform the single best classifier. If so, what form of ensemble learning system (also known as multiple classifier learning…
Jonathan S. Kent, David H. Menager
For many years, researchers and practitioners have designed and employed rule-based classification systems. In contrast with other machine learning classifiers, rule-based systems are useful, not only because of their predictive power, but also because they often produce knowledge structures, or rules, humans readily…
Dimitris Bertsimas, Vassilios Digalakis
Owing to their inherently interpretable structure, decision trees are commonly used in applications where interpretability is essential. Recent work has focused on improving various aspects of decision trees, including their predictive power and robustness; however, their instability, albeit well-documented, has been…
Hendrik Blockeel, Laurens Devos, Benoît Frénay, Géraldin Nanfack + 1 more
'Siegfried Nijssen'] This article provides a birds-eye view on the role of decision trees in machine learning and data science over roughly four decades. It sketches the evolution of decision tree research over the years, describes the broader context in which the research is situated, and summarizes strengths and…
Thomas A. Lake, Brit B. Laginhas, Brennen T. Farrell, Ross K. Meentemeyer + 1 more
Accurate and up-to-date catalogs of urban tree populations are crucial for quantifying ecosystem services and enhancing the quality of life in cities. However, identifying and mapping tree species cost-effectively remains a significant challenge. Remote sensing is an active area of research where scientists are…
Tuomas Aakala, Juha Heikkinen
Dead wood quality is recorded as a biodiversity indicator and in estimating forest ecosystem carbon storage, using decay classification systems. In large-scale national forest inventories (NFIs), these systems are typically slightly different among countries, but harmonizing them would allows analyses over much broader…
Pablo del Moral, Sławomir Nowaczyk, Anita Sant’Anna, Sepideh Pashami
Using hierarchies of classes is one of the standard methods to solve multi-class classification problems. In the literature, selecting the right hierarchy is considered to play a key role in improving classification performance. Although different methods have been proposed, there is still a lack of understanding of…
James G C Ball, Sadiq Jaffer, Anthony Laybros, Colin Prieur + 5 more
To understand how tropical rainforests will adapt to climate change and the extent to which their diversity imparts resilience, precise, taxonomically informed monitoring of individual trees is required. However, the density, diversity and complexity of tropical rainforests present considerable challenges to remote…
Prashanth Athri, Vidhya Murali, Pradyumna Y Muralidhar, Cassandra Königs + 4 more
- 1. Department of Computer Science and Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham, Bengaluru, India - 2. PES Center for Pattern Recognition, Department of Computer Science and Engineering, PES University, Bengaluru, India - 3. Bioinformatics and Medical Informatics, Bielefeld University…
Esteban Bertsch Aguilar, Sebastián Suñer Sánchez, Silvana Pinheiro, William J. Zamora Ramírez
- 1. 1. CBio3 Laboratory, School of Chemistry, University of Costa Rica, San Pedro, San José, Costa Rica - 2. 2. Laboratory of Computational Toxicology and Artificial Intelligence (LaToxCIA), Biological Testing Laboratory (LEBi), University of Costa Rica, San Pedro, San José, Costa Rica - 3. 3. Advanced Computing Lab…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European Chemicals Agency (ECHA) and US Environmental Protection Agency (EPA) have listed approximately 800k chemicals that must be further investigated for their potential environmental and/or human health risk. A significant number of these chemicals have large enough global volumes of consumption (e.g.…
Tianjian Qin, Koen van Benthem, Luis Valente, Rampal Etienne
Reconstructing the forces that shaped macroevolutionary histories from extant phylogenies is fundamentally challenging: richly parameterized diversification models are often only weakly identifiable; different evolutionary mechanisms can yield nearly indistinguishable tree shapes. Here we use a model with evolutionary…
Itamar Borges Jr, Júlio César Duarte, Romulo Dias da Rocha
We decomposed density functional theory charge densities of 53 nitroaromatic molecules into atom-centered electric multipoles using the distributed multipole analysis that provides a detailed picture of the molecular electronic structure. Three electric multipoles, ∑▒〖Q_0 (NO_2)〗 (the charge of the nitro groups)…
Jonas Schaub, Julian Zander, Achim Zielesny, Christoph Steinbeck
The concept of molecular scaffolds as defining core structures of organic molecules is utilised in many areas of chemistry and cheminformatics, e.g. drug design, chemical classification, or the analysis of high-throughput screening data. Here, we present Scaffold Generator, a comprehensive open library for the…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European Chemicals Agency (ECHA) and US Environmental Protection Agency (EPA) have listed approximately 800k chemicals that must be further investigated for their potential environmental and/or human health risk. A significant number of these chemicals have large enough global volumes of consumption (e.g.…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European and US chemical agencies have listed approximately 800k chemicals where knowledge on potential risks to human health and the environment are lacking. Filling these data gaps experimentally is impossible so in-silico approaches and prediction are essential. Many existing models are however limited by…