28 papers · ranked by Valyu relevance
Gérard Biau, Erwan Scornet
The random forest algorithm, proposed by L. Breiman in 2001, has been extremely successful as a general-purpose classification and regression method. The approach, which combines several randomized decision trees and aggregates their predictions by averaging, has shown excellent performance in settings where the number…
Li Ma, Suohai Fan
Background The random forests algorithm is a type of classifier with prominent universality, a wide application range, and robustness for avoiding overfitting. But there are still some drawbacks to random forests. Therefore, to improve the performance of random forests, this paper seeks to improve imbalanced data…
Hemant Ishwaran, James D Malley
Background Using a collection of different terminal nodesize constructed random forests, each generating a synthetic feature, a synthetic random forest is defined as a kind of hyperforest, calculated using the new input synthetic features, along with the original features. Results Using a large collection of regression…
Wiem Elghazel, Kamal Medjaher, Noureddine Zerhouni, Jacques M. Bahi + 3 more
'Ahmad Farhat' 'Christophe Guyeux' 'Mourad Hakem'] In this paper, random forests are proposed for operating devices diagnostics in the presence of a variable number of features. In various contexts, like large or difficult-to-access monitored areas, wired sensor networks providing features to achieve diagnostics are…
Roman Hornung
The diversity forest algorithm is an alternative candidate node split sampling scheme that makes innovative complex split procedures in random forests possible. While conventional univariable, binary splitting suffices for obtaining strong predictive performance, new complex split procedures can help tackling…
Nicholas Waltz
Despite their performance and widespread use, little is known about the theory of Random Forests. A major unanswered question is whether, or when, the Random Forest algorithm is consistent. The literature explores various variants of the classic Random Forest algorithm to address this question and known short-comings…
Maria C Mariani, Osei K Tweneboah, Md Al Masum Bhuiyan
This work analyses the diagnosis and prognosis of cancer and heart disease data using five Machine Learning (ML) algorithms. We compare the predictive ability of all the ML algorithms to breast cancer and heart disease. The important variables that causes cancer and heart disease are also studied. We predict the test…
Gunther Schauberger, Stefanie J. Klug, Moritz Berger
Background Conditional logistic regression trees have been proposed as a flexible alternative to the standard method of conditional logistic regression for the analysis of matched case-control studies. While they allow to avoid the strict assumption of linearity and automatically incorporate interactions, conditional…
Roxane Duroux, Erwan Scornet
Random forests are ensemble learning methods introduced by Breiman (2001) that operate by averaging several decision trees built on a randomly selected subspace of the data set. Despite their widespread use in practice, the respective roles of the different mechanisms at work in Breiman's forests are not yet fully…
Hao Wu, Qiaomei Wang, Kunjian Yu, Xiaotao Hu + 2 more
In order to remedy the current problem of having been buffeted by competing requirements for both protection sensitivity and quick reaction of High Voltage Direct Current (HVDC) transmission lines simultaneously, a new intelligent fault identification method based on Random Forests (RF) for HVDC transmission lines is…
Björn-Hergen Laabs von Holt, Ana Westenberger, Inke R. König
In life sciences random forests are often used to train predictive models. However, gaining any explanatory insight into the mechanics leading to a specific outcome is rather complex, which impedes the implementation of random forests into clinical practice. By simplifying a complex ensemble of decision trees to a…
Jianyuan Sun, Guoqiang Zhong, Junyu Dong, Yajuan Cai
Random forests are a type of ensemble method which makes predictions by combining the results of several independent trees. However, the theory of random forests has long been outpaced by their application. In this paper, we propose a novel random forests algorithm based on cooperative game theory. Banzhaf power index…
Iakovidis Isidoros, Nicola Arcozzi
Random forest methods belong to the class of non-parametric machine learning algorithms. They were first introduced in 2001 by Breiman and they perform with accuracy in high dimensional settings. In this article, we consider a simplified kernel-based random forest algorithm called simplified directional KeRF (Kernel…
Yi Wang, Yi Li, Weilin Pu, Kathryn Wen + 3 more
'Momiao Xiong' 'Li Jin'] Efficiency, memory consumption, and robustness are common problems with many popular methods for data analysis. As a solution, we present Random Bits Forest (RBF), a classification and regression algorithm that integrates neural networks (for depth), boosting (for width), and random forests…
Roozbeh Valavi, Jane Elith, José J. Lahoz-Monfort, Gurutzeta Guillera-Arroita
The Random Forest (RF) algorithm is an ensemble of classification or regression trees, and is a widely used and high-performing machine learning technique. It is increasingly used for species distribution modelling (SDM). Many researchers use implementations of RF in the R programming language with default parameters…
Arash Bayat, Piotr Szul, Aidan R. O’Brien, Robert Dunne + 4 more
The demands on machine learning methods to cater for ultra high dimensional datasets, datasets with millions of features, have been increasing in domains like life sciences and the Internet of Things (IoT). While Random Forests are suitable for “wide” datasets, current implementations such as Google’s PLANET lack the…
Krzysztof Gajowniczek, Iga Grzegorczyk, Tomasz Ząbkowski
The literature indicates that 90% of clinical alarms in intensive care units might be false. This high percentage negatively impacts both patients and clinical staff. In patients, false alarms significantly increase stress levels, which is especially dangerous for cardiac patients. In clinical staff, alarm overload…
Samir Rachid Zaim, Colleen Kenost, Joanne Berghout, Wesley Chiu + 3 more
In this era of data science-driven bioinformatics, machine learning research has focused on feature selection as users want more interpretation and post-hoc analyses for biomarker detection. However, when there are more features (i.e., transcript) than samples (i.e., mice or human samples) in a study, this poses major…
Lillian Oluoch, László Stachó, László Viharos, Andor Viharos + 1 more
To overcome well-known difficulties in establishing reliable models based on large data sets, the Random Forest Regression (RFR) method is applied to study economical breeding and milk production of dairy cows. As for the features of RFR, there are several positive experiences in various areas of applications…
Maziyar Baran Pouyan, Dennis Kostka
Genome-wide transcriptome sequencing applied to single cells (scRNA-seq) is rapidly becoming an assay of choice across many fields of biological and biomedical research. Scientific objectives often revolve around discovery or characterization of types or sub-types of cells, and therefore obtaining accurate cell–cell…
Tung Dang, Hirohisa Kishino
Random forest (RF) captures complex feature patterns that differentiate groups of samples and is rapidly being adopted in microbiome studies. However, a major challenge is the high dimensionality of microbiome datasets. They include thousands of species or molecular functions of particular biological interest. This…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Prashanth Athri, Vidhya Murali, Pradyumna Y Muralidhar, Cassandra Königs + 4 more
- 1. Department of Computer Science and Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham, Bengaluru, India - 2. PES Center for Pattern Recognition, Department of Computer Science and Engineering, PES University, Bengaluru, India - 3. Bioinformatics and Medical Informatics, Bielefeld University…
Authors not listed
The dual imperative of mitigating carbon emissions and maximizing hydrocarbon recovery has amplified global interest in carbon capture, utilization, and storage (CCUS) technologies. These integrated processes hold significant promise for achieving net-zero targets while extending the productive life of mature oil…
Moayad Alnammi, Shengchao Liu, Spencer S Ericksen, Gene E Ananiev + 6 more
Traditional small molecule drug discovery is a time consuming and costly endeavor. High-throughput chemical screening can only assess a tiny fraction of drug-like chemical space. The strong predictive power of modern machine learning methods for virtual chemical screening enables training models on known active and…
Devi Ganapathi, Wunmi Akinlemibola, Antonio Baclig, Emily Penn + 1 more
Quinones and hydroquinones are small organic molecules with numerous applications: battery electrolytes, pharmaceuticals, sensors, to name a few. An understanding of their fundamental properties, such as melting points, is essential to incorporate these compounds into relevant technologies. In this study, two different…
Authors not listed
Crystal structure prediction (CSP) is a valuable computational technique used to anticipate the likely crystal structures of a compound of interest. These methods have been proven useful in research and development of pharmaceutical solid forms and in guiding the discovery of materials with targeted properties. Despite…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…