27 papers · ranked by Valyu relevance
Gérard Biau, Erwan Scornet
The random forest algorithm, proposed by L. Breiman in 2001, has been extremely successful as a general-purpose classification and regression method. The approach, which combines several randomized decision trees and aggregates their predictions by averaging, has shown excellent performance in settings where the number…
Carolin Strobl, Anne-Laure Boulesteix, Thomas Kneib, Thomas Augustin + 1 more
'Achim Zeileis'] Background Random forests are becoming increasingly popular in many scientific fields because they can cope with "small n large p" problems, complex interactions and even highly correlated predictor variables. Their variable importance measures have recently been suggested as screening tools for, e.g.…
Hemant Ishwaran, James D Malley
Background Using a collection of different terminal nodesize constructed random forests, each generating a synthetic feature, a synthetic random forest is defined as a kind of hyperforest, calculated using the new input synthetic features, along with the original features. Results Using a large collection of regression…
Wiem Elghazel, Kamal Medjaher, Noureddine Zerhouni, Jacques M. Bahi + 3 more
'Ahmad Farhat' 'Christophe Guyeux' 'Mourad Hakem'] In this paper, random forests are proposed for operating devices diagnostics in the presence of a variable number of features. In various contexts, like large or difficult-to-access monitored areas, wired sensor networks providing features to achieve diagnostics are…
Mohammad Savargiv, Behrooz Masoumi, Mohammad Reza Keyvanpour
The goal of aggregating the base classifiers is to achieve an aggregated classifier that has a higher resolution than individual classifiers. Random forest is one of the types of ensemble learning methods that have been considered more than other ensemble learning methods due to its simple structure, ease of…
Roman Hornung
The diversity forest algorithm is an alternative candidate node split sampling scheme that makes innovative complex split procedures in random forests possible. While conventional univariable, binary splitting suffices for obtaining strong predictive performance, new complex split procedures can help tackling…
Helen L. Smith, Patrick J. Biggs, Nigel P. French, Adam N. H. Smith + 1 more
To date, there remains no satisfactory solution for absent levels in random forest models. Absent levels are levels of a predictor variable encountered during prediction for which no explicit rule exists. Imposing an order on nominal predictors allows absent levels to be integrated and used for prediction. The ordering…
Gunther Schauberger, Stefanie J. Klug, Moritz Berger
Background Conditional logistic regression trees have been proposed as a flexible alternative to the standard method of conditional logistic regression for the analysis of matched case-control studies. While they allow to avoid the strict assumption of linearity and automatically incorporate interactions, conditional…
Cole Brokamp, M. Bhaskara Rao, Patrick Ryan, Roman Jandarov
The infinitesimal jackknife (IJ) has recently been applied to the random forest to estimate its prediction variance. These theorems were verified under a traditional random forest framework which uses classification and regression trees (CART) and bootstrap resampling. However, random forests using conditional…
Björn-Hergen Laabs von Holt, Ana Westenberger, Inke R. König
In life sciences random forests are often used to train predictive models. However, gaining any explanatory insight into the mechanics leading to a specific outcome is rather complex, which impedes the implementation of random forests into clinical practice. By simplifying a complex ensemble of decision trees to a…
Delilah Donick, Sandro Claudio Lera
Conventionally, random forests are built from “greedy” decision trees which each consider only one split at a time during their construction. The sub-optimality of greedy implementation has been well-known, yet mainstream adoption of more sophisticated tree building algorithms has been lacking. We examine under what…
Roxane Duroux, Erwan Scornet
Random forests are ensemble learning methods introduced by Breiman (2001) that operate by averaging several decision trees built on a randomly selected subspace of the data set. Despite their widespread use in practice, the respective roles of the different mechanisms at work in Breiman's forests are not yet fully…
Misha Denil, David S. Matheson, Nando de Freitas
Random forests are a class of ensemble method whose base learners are a collection of randomized tree predictors, which are combined through averaging. The original random forests framework described in Breiman (2001) has been extremely influential (Svetnik et al., 2003; Prasad et al., 2006; Cutler et al., 2007…
Iakovidis Isidoros, Nicola Arcozzi
Random forest methods belong to the class of non-parametric machine learning algorithms. They were first introduced in 2001 by Breiman and they perform with accuracy in high dimensional settings. In this article, we consider a simplified kernel-based random forest algorithm called simplified directional KeRF (Kernel…
Jianyuan Sun, Guoqiang Zhong, Junyu Dong, Yajuan Cai
Random forests are a type of ensemble method which makes predictions by combining the results of several independent trees. However, the theory of random forests has long been outpaced by their application. In this paper, we propose a novel random forests algorithm based on cooperative game theory. Banzhaf power index…
Arash Bayat, Piotr Szul, Aidan R. O’Brien, Robert Dunne + 4 more
The demands on machine learning methods to cater for ultra high dimensional datasets, datasets with millions of features, have been increasing in domains like life sciences and the Internet of Things (IoT). While Random Forests are suitable for “wide” datasets, current implementations such as Google’s PLANET lack the…
Yi Wang, Yi Li, Weilin Pu, Kathryn Wen + 3 more
'Momiao Xiong' 'Li Jin'] Efficiency, memory consumption, and robustness are common problems with many popular methods for data analysis. As a solution, we present Random Bits Forest (RBF), a classification and regression algorithm that integrates neural networks (for depth), boosting (for width), and random forests…
Tyler Kolisnik, Faeze Keshavarz-Rahaghi, Rachel Purcell, Adam Smith + 1 more
Random Forest models are widely used in genomic data analysis and can offer insights into complex biological mechanisms, particularly where features influence the target in interactive, non-linear, or non-additive ways. Currently, some of the most efficient random forest methods, in terms of computational speed, are…
Xiangkui Jiang, Chang-an Wu, Huaping Guo
A forest is an ensemble with decision trees as members. This paper proposes a novel strategy to pruning forest to enhance ensemble generalization ability and reduce ensemble size. Unlike conventional ensemble pruning approaches, the proposed method tries to evaluate the importance of branches of trees with respect to…
Maziyar Baran Pouyan, Dennis Kostka
Genome-wide transcriptome sequencing applied to single cells (scRNA-seq) is rapidly becoming an assay of choice across many fields of biological and biomedical research. Scientific objectives often revolve around discovery or characterization of types or sub-types of cells, and therefore obtaining accurate cell–cell…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Authors not listed
The dual imperative of mitigating carbon emissions and maximizing hydrocarbon recovery has amplified global interest in carbon capture, utilization, and storage (CCUS) technologies. These integrated processes hold significant promise for achieving net-zero targets while extending the productive life of mature oil…
Moayad Alnammi, Shengchao Liu, Spencer S Ericksen, Gene E Ananiev + 6 more
Traditional small molecule drug discovery is a time consuming and costly endeavor. High-throughput chemical screening can only assess a tiny fraction of drug-like chemical space. The strong predictive power of modern machine learning methods for virtual chemical screening enables training models on known active and…
Authors not listed
We present a transferable, interpretable, and modular machine-learning framework that enhances the accuracy of density functional theory (DFT) reaction energies using physically meaningful energy-decomposition descriptors. Reaction energies computed at the DFT level with standard basis sets are first decomposed into…
Aarav Arora, Igor Tsigelny, Valentina Kouznetsova
Purpose Laryngeal cancer (LC) is the most common head and neck cancer, which often goes undiagnosed due to the expensiveness and inaccessible nature of current diagnosis methods. Many recent studies have shown that microRNAs (miRNAs) are crucial biomarkers for a variety of cancers. Methods In this study, we create a…
Authors not listed
Phase equilibrium calculations are crucial in chemical engineering design and optimization processes. The PC-SAFT equation of state (EoS) can precisely calculate phase equilibrium, but is relatively complex and computationally intensive. Surrogate models are mathematically simple models that map or regress the…
Authors not listed
Background: Janus Kinase 2 (JAK2) is a key kinase in cellular signal transduction. Its abnormal activation is closely related to various myeloproliferative neoplasms and inflammatory diseases. Developing selective JAK2 inhibitors is an important direction in drug discovery. Accurate prediction of compound inhibitory…