28 papers · ranked by Valyu relevance
Zhi-Hua Zhou, Ji Feng
Current deep-learning models are mostly built upon neural networks, i.e. multiple layers of parameterized differentiable non-linear modules that can be trained by backpropagation. In this paper, we explore the possibility of building deep models based on non-differentiable modules such as decision trees. After a…
Palak Mahajan, Shahadat Uddin, Farshid Hajati, Mohammad Ali Moni + 1 more
'Joaquim Carreras'] Machine learning models are used to create and enhance various disease prediction frameworks. Ensemble learning is a machine learning technique that combines multiple classifiers to improve performance by making more accurate predictions than a single classifier. Although numerous studies have…
Ahmed Ali Mohamed Warad, Khaled Wassif, Nagy Ramadan Darwish
Based on the benefits of different ensemble methods, such as bagging and boosting, which have been studied and adopted extensively in research and practice, where bagging and boosting focus more on reducing variance and bias, this paper presented an optimization ensemble learning-based model for a large pipe failure…
Halit Karalar, Ceyhun Kapucu, Hüseyin Gürüler
Predicting students at risk of academic failure is valuable for higher education institutions to improve student performance. During the pandemic, with the transition to compulsory distance learning in higher education, it has become even more important to identify these students and make instructional interventions to…
Cesar Alfaro, Javier Gomez, Javier M. Moguerza, Javier Castillo + 2 more
'Jose I. Martinez' 'Sotiris Kotsiantis'] Typical applications of wireless sensor networks (WSN), such as in Industry 4.0 and smart cities, involves acquiring and processing large amounts of data in federated systems. Important challenges arise for machine learning algorithms in this scenario, such as reducing energy…
Roberto Aldave, Jean‐Pierre Dussault
The motivation of this work is to improve the performance of standard stacking approaches or ensembles, which are composed of simple, heterogeneous base models, through the integration of the generation and selection stages for regression problems. We propose two extensions to the standard stacking approach. In the…
Xiaoye Mo, Xia Jiang
Ubiquitination-site prediction is an important task because ubiquitination is a critical regulatory function for many biological processes such as proteasome degradation, DNA repair and transcription, signal transduction, endocytoses, and sorting. However, the highly dynamic and reversible nature of ubiquitination…
Kuo-Wei Hsu
Inspired by the group decision making process, ensembles or combinations of classifiers have been found favorable in a wide variety of application domains. Some researchers propose to use the mixture of two different types of classification algorithms to create a hybrid ensemble. Why does such an ensemble work? The…
Shikun Chen, Wenlong Zheng, Syed Nisar Hussain Bukhari,
Ensemble regression methods are widely used to improve prediction accuracy by combining multiple regression models, especially when dealing with continuous numerical targets. However, most ensemble voting regressors use equal weights for each base model’s predictions, which can limit their effectiveness, particularly…
Ling Liu, Wenqi Wei, Ka-Ho Chow, Margaret L. Loper + 3 more
'Stacey Truex' 'Yanzhao Wu'] Abstract—Ensemble learning is a methodology that integrates multiple DNN learners for improving prediction performance of individual learners. Diversity is greater when the errors of the ensemble prediction is more uniformly distributed. Greater diversity is highly correlated with the…
Wenjing Li, Randy Paffenroth, David Berthiaume
Ensemble learning is a process by which multiple base learners are strategically generated and combined into one composite learner. There are two features that are essential to an ensemble's performance, the individual accuracies of the component learners and the overall diversity in the ensemble. The right balance of…
Vanda M. Lourenço, Joseph O. Ogutu, Rui A.P. Rodrigues, Hans-Peter Piepho
The accurate prediction of genomic breeding values is central to genomic selection in both plant and animal breeding studies. Genomic prediction involves the use of thousands of molecular markers spanning the entire genome and therefore requires methods able to efficiently handle high dimensional data. Not…
Arjun Pakrashi, Brian Mac Namee
This paper introduces a new perspective on multi-class ensemble classification that considers training an ensemble as a state estimation problem. The new perspective considers the final ensemble classifier model as a static state, which can be estimated using a Kalman filter that combines noisy estimates made by…
Jay Devine, Helen K. Kurki, Jonathan R. Epp, Paula N. Gonzalez + 2 more
Classification is a fundamental task in biology used to assign members to a class. While linear discriminant functions have long been effective, advances in phenotypic data collection are yielding increasingly high-dimensional datasets with more classes, unequal class covariances, and non-linear distributions. Numerous…
Michael Onyema Edeh, Surjeet Dalal, Imed Ben Dhaou, Charles Chuka Agubosim + 3 more
Machine learning algorithms are excellent techniques to develop prediction models to enhance response and efficiency in the health sector. It is the greatest approach to avoid the spread of hepatitis C, especially injecting drugs, is to avoid these behaviors. Treatments for hepatitis C can cure most patients within 8…
Moshe Sipper
> Abstract. We present Classy Ensemble, a novel ensemble-generation algorithm for classification tasks, which aggregates models through a weighted combination of per-class accuracy. Tested over 153 machine learning datasets we demonstrate that Classy Ensemble outperforms two other well-known aggregation…
Pierre‐Alexandre Mattei, Damien Garreau
Ensemble methods combine the predictions of several base models. We study whether or not including more models always improves their average performance. This question depends on the kind of ensemble considered, as well as the predictive metric chosen. We focus on situations where all members of the ensemble are a…
Yeming Wen, Dustin Tran, Jimmy Ba
Ensembles, where multiple neural networks are trained individually and their predictions are averaged, have been shown to be widely successful for improving both the accuracy and predictive uncertainty of single neural networks. However, an ensemble's cost for both training and testing increases linearly with the…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Abdul Mueed Hafiz, G. Mohiuddin Bhat
- Traditional machine learning approaches may fail to perform satisfactorily when dealing with complex data. In this context, the importance of data mining evolves w.r.t. building an efficient knowledge discovery and mining framework. Ensemble learning is aimed at integration of fusion, modeling and mining of data into…
Maya Ramchandran, Prasad Patil, Giovanni Parmigiani
Multi-study learning uses multiple training studies, separately trains classifiers on individual studies, and then forms ensembles with weights rewarding members with better cross-study prediction ability. This article considers novel weighting approaches for constructing tree-based ensemble learners in this setting.…
Moayad Alnammi, Shengchao Liu, Spencer S Ericksen, Gene E Ananiev + 6 more
Traditional small molecule drug discovery is a time consuming and costly endeavor. High-throughput chemical screening can only assess a tiny fraction of drug-like chemical space. The strong predictive power of modern machine learning methods for virtual chemical screening enables training models on known active and…
H.M.Fazlul Haque, Fariha Arifin, Sheikh Adilina, Muhammod Rafsanjani + 1 more
The information of a cell is primarily contained in Deoxyribonucleic Acid (DNA). There is a flow of information of DNA to protein sequences via Ribonucleic acids (RNA) through transcription and translation. These entities are vital for the genetic process. Recent developments in epigenetic also show the importance of…
Authors not listed
In molecular machine learning, the choice of the representation of molecules can have a significant impact on model performance. However, understanding the root causes of these performance differences often proves challenging. One promising approach to explore model behavior is representational alignment, which…
Esther Heid, Charles J. McGill, Florence H. Vermeire, William H. Green
Characterizing uncertainty in machine learning models has recently gained interest in the context of machine learning reliability, robustness, safety, and active learning. Here, we separate the total uncertainty into contributions from noise in the data (aleatoric) and shortcomings of the model (epistemic), further…
Authors not listed
Accurate prediction of redox potentials of iron (Fe) complexes, in tandem with uncertainty quantification, is essential to advance technologies related to electro-deposition and energy storage by enabling reliable modeling, guiding experimental design, and improving the efficiency of material discovery. Since…
Paul Francoeur, Daniel Penaherrera, David Koes
The immense size of chemical space, the relative scarcity of high quality data, and the cost of running experiments to accurately measure molecular properties makes active learning (AL) an attractive approach to efficiently explore the space and train high-quality models for molecular property prediction. While AL is…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…