28 papers · ranked by Valyu relevance
Hendrik Blockeel, Laurens Devos, Benoît Frénay, Géraldin Nanfack + 1 more
'Siegfried Nijssen'] This article provides a birds-eye view on the role of decision trees in machine learning and data science over roughly four decades. It sketches the evolution of decision tree research over the years, describes the broader context in which the research is situated, and summarizes strengths and…
M. Nabipour, P. Nayyeri, H. Jabani, A. Mosavi + 2 more
The prediction of stock groups values has always been attractive and challenging for shareholders due to its inherent dynamics, non-linearity, and complex nature. This paper concentrates on the future prediction of stock market groups. Four groups named diversified financials, petroleum, non-metallic minerals, and…
Zebin Yang, Agus Sudjianto, Xiaoming Li, Aijun Zhang
—Tree ensemble models like random forests and gradient boosting machines are widely used in machine learning due to their excellent predictive performance. However, a highperformance ensemble consisting of a large number of decision trees lacks sufficient transparency and explainability. In this paper, we demonstrate…
Authors not listed
Phase equilibrium calculations are crucial in chemical engineering design and optimization processes. The PC-SAFT equation of state (EoS) can precisely calculate phase equilibrium, but is relatively complex and computationally intensive. Surrogate models are mathematically simple models that map or regress the…
Mohammed Ghazwani, M. Yasmin Begum
This work presents the results of using tree-based models, including Gradient Boosting, Extra Trees, and Random Forest, to model the solubility of hyoscine drug and solvent density based on pressure and temperature as inputs. The models were trained on a dataset of hyoscine drug with known solubility and density…
Patrick J. Miller, Gitta H. Lubke, Daniel B. McArtor, C. S. Bergeman
This research was based upon work supported by the National Science Foundation Graduate Research Fellowship Program under grant number 1313583. The second author is supported by NIDA R37 DA-018673. The fourth author is supported by a grant from the National Institute of Aging (1 R01 AG023571-A1-01). The computational…
Saadin Oyucu, Onur Polat, Muammer Türkoğlu, Hüseyin Polat + 3 more
'Ahmet Aksöz' 'Mehmet Tevfik Ağdaş' 'Naveen Chilamkurti'] Supervisory Control and Data Acquisition (SCADA) systems play a crucial role in overseeing and controlling renewable energy sources like solar, wind, hydro, and geothermal resources. Nevertheless, with the expansion of conventional SCADA network infrastructures…
Gitesh Dawer, Yangzi Guo, Adrian Barbu
—Tree ensembles are flexible predictive models that can capture relevant variables and to some extent their interactions in a compact and interpretable manner. Most algorithms for obtaining tree ensembles are based on versions of boosting or Random Forest. Previous work showed that boosting algorithms exhibit a cyclic…
D. P. P. Meddage, I. U. Ekanayake, Sumudu Herath, R. Gobirahavan + 3 more
'Nitin Muttil' 'Upaka Rathnayake' 'Roberto Teti'] Predicting the bulk-average velocity (UB) in open channels with rigid vegetation is complicated due to the non-linear nature of the parameters. Despite their higher accuracy, existing regression models fail to highlight the feature importance or causality of the…
Hadrien Bride, Zhé Hóu, Jie Dong, Jin Song Dong + 1 more
'Seyed Mohammad Mirjalili'] This paper introduces a new classification tool named Silas, which is built to provide a more transparent and dependable data analytics service. A focus of Silas is on providing a formal foundation of decision trees in order to support logical analysis and verification of learned prediction…
Liangyuan Hu, Lihua Li, Paul B. Tchounwou
Tree-based machine learning methods have gained traction in the statistical and data science fields. They have been shown to provide better solutions to various research questions than traditional analysis approaches. To encourage the uptake of tree-based methods in health research, we review the methodological…
Leo L. Duan, John Clancy, Rhonda D. Szczesniak
We propose a novel "tree-averaging" model that utilizes the ensemble of classification and regression trees (CART). Each constituent tree is estimated with a subset of similar data. We treat this grouping of subsets as Bayesian ensemble trees (BET) and model them as an infinite mixture Dirichlet process. We show that…
Xiaoye Mo, Xia Jiang
Ubiquitination-site prediction is an important task because ubiquitination is a critical regulatory function for many biological processes such as proteasome degradation, DNA repair and transcription, signal transduction, endocytoses, and sorting. However, the highly dynamic and reversible nature of ubiquitination…
Alicia Curth, Alan Jeffares, Mihaela van der Schaar
Self-Regularizing Adaptive Smoothers Authors: ['Alicia Curth' 'Alan Jeffares' 'Mihaela van der Schaar'] Despite their remarkable effectiveness and broad application, the drivers of success underlying ensembles of trees (especially random forests and gradient boosting) are still not fully understood. In this paper, we…
Andrei V. Konstantinov, Lev V. Utkin
The gradient boosting machine is a powerful ensemble-based machine learning method for solving regression problems. However, one of the difficulties of its using is a possible discontinuity of the regression function, which arises when regions of training data are not densely covered by training points. In order to…
Maya Ramchandran, Prasad Patil, Giovanni Parmigiani
Multi-study learning uses multiple training studies, separately trains classifiers on individual studies, and then forms ensembles with weights rewarding members with better cross-study prediction ability. This article considers novel weighting approaches for constructing tree-based ensemble learners in this setting.…
Sebastian Spänig, Alexander Michel, Dominik Heider
Owing to the rising levels of multi-resistant pathogens, antimicrobial peptides, an alternative strategy to classic antibiotics, got more attention. A crucial part is thereby the costly identification and validation. With the ever-growing amount of annotated peptides, researchers employed artificial intelligence to…
Robert C. Edgar
Phylogenetic tree confidence is often estimated from a multiple sequence alignment (MSA) using the Felsenstein bootstrap heuristic. However, this does not account for systematic errors in the MSA, which may cause substantial bias to the inferred phylogeny. Here, I describe the MSA ensemble bootstrap, a new procedure…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Prashanth Athri, Vidhya Murali, Pradyumna Y Muralidhar, Cassandra Königs + 4 more
- 1. Department of Computer Science and Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham, Bengaluru, India - 2. PES Center for Pattern Recognition, Department of Computer Science and Engineering, PES University, Bengaluru, India - 3. Bioinformatics and Medical Informatics, Bielefeld University…
Ahmed Ali Mohamed Warad, Khaled Wassif, Nagy Ramadan Darwish
Based on the benefits of different ensemble methods, such as bagging and boosting, which have been studied and adopted extensively in research and practice, where bagging and boosting focus more on reducing variance and bias, this paper presented an optimization ensemble learning-based model for a large pipe failure…
Michał Gostkowski, Krzysztof Gajowniczek
Due to various regulations (e.g., the Basel III Accord), banks need to keep a specified amount of capital to reduce the impact of their insolvency. This equity can be calculated using, e.g., the Internal Rating Approach, enabling institutions to develop their own statistical models. In this regard, one of the most…
Authors not listed
Accurate prediction of redox potentials of iron (Fe) complexes, in tandem with uncertainty quantification, is essential to advance technologies related to electro-deposition and energy storage by enabling reliable modeling, guiding experimental design, and improving the efficiency of material discovery. Since…
Zijie Zhao, Stephen Dorn, Yuchang Wu, Xiaoyu Yang + 2 more
Ensemble learning has been increasingly popular for boosting the predictive power of polygenic risk scores (PRS), with almost every recent multi-ancestry PRS approach employing ensemble learning as a final step. Existing ensemble approaches rely on individual-level data for model training, which severely limits their…
Authors not listed
Machine learning holds significant promise for accelerating biomarker discovery in clinical proteomics, yet its real-world impact remains limited by widespread methodological pitfalls and unrealistic expectations. In this perspective, we critically examine the integration of machine learning into clinical proteomics…
Aarav Arora, Igor Tsigelny, Valentina Kouznetsova
Purpose Laryngeal cancer (LC) is the most common head and neck cancer, which often goes undiagnosed due to the expensiveness and inaccessible nature of current diagnosis methods. Many recent studies have shown that microRNAs (miRNAs) are crucial biomarkers for a variety of cancers. Methods In this study, we create a…
Authors not listed
Metal hydrides play a pivotal role in a wide range of applications, including hydrogen storage, compression, heat management, and catalysis, making them a central focus of interdisciplinary research spanning chemistry, materials science, and engineering. The performance of the metal hydride based systems is strongly…
Authors not listed
A protocol for generating potential energy surfaces and performing photoinduced nonadiabatic multidimensional wave packet propagation is presented. The workflow starts with the parameterization of a linear vibronic coupling (LVC) Hamiltonian using the BSE@GW approach. In a second step, the LVC model is used as input…