26 papers · ranked by Valyu relevance
Roman Hornung
The diversity forest algorithm is an alternative candidate node split sampling scheme that makes innovative complex split procedures in random forests possible. While conventional univariable, binary splitting suffices for obtaining strong predictive performance, new complex split procedures can help tackling…
Joel Therrien, Jiguo Cao
Random forests are a sensible non-parametric model to predict competing risks data according to some covariates. However, there are currently no packages that can adequately handle large datasets (n > 100, 000). We introduce a new R package, largeRCRF, using the random competing risks forest theory developed by…
Helen L. Smith, Patrick J. Biggs, Nigel P. French, Adam N. H. Smith + 1 more
To date, there remains no satisfactory solution for absent levels in random forest models. Absent levels are levels of a predictor variable encountered during prediction for which no explicit rule exists. Imposing an order on nominal predictors allows absent levels to be integrated and used for prediction. The ordering…
Nicholas Waltz
Despite their performance and widespread use, little is known about the theory of Random Forests. A major unanswered question is whether, or when, the Random Forest algorithm is consistent. The literature explores various variants of the classic Random Forest algorithm to address this question and known short-comings…
Brian Liu, Rahul Mazumder
Forests Authors: ['Brian Liu' 'Rahul Mazumder'] We study the often overlooked phenomenon, first noted in Breiman (2001), that random forests appear to reduce bias compared to bagging. Motivated by an interesting paper by Mentch and Zhou (2020), where the authors argue that random forests reduce effective degrees of…
Björn-Hergen Laabs von Holt, Ana Westenberger, Inke R. König
In life sciences random forests are often used to train predictive models. However, gaining any explanatory insight into the mechanics leading to a specific outcome is rather complex, which impedes the implementation of random forests into clinical practice. By simplifying a complex ensemble of decision trees to a…
Matias D. Cattaneo, Jason M. Klusowski, William G. Underwood
Random forests are popular methods for regression and classification analysis, and many different variants have been proposed in recent years. One interesting example is the Mondrian random forest, in which the underlying constituent trees are constructed via a Mondrian process. We give precise bias and variance…
Vera Ignatenko, Anton Surkov, Sergei Koltcov, Bilal Alatas
The random forest algorithm is one of the most popular and commonly used algorithms for classification and regression tasks. It combines the output of multiple decision trees to form a single result. Random forest algorithms demonstrate the highest accuracy on tabular data compared to other algorithms in various…
Gunther Schauberger, Stefanie J. Klug, Moritz Berger
Background Conditional logistic regression trees have been proposed as a flexible alternative to the standard method of conditional logistic regression for the analysis of matched case-control studies. While they allow to avoid the strict assumption of linearity and automatically incorporate interactions, conditional…
Iakovidis Isidoros, Nicola Arcozzi
Random forest methods belong to the class of non-parametric machine learning algorithms. They were first introduced in 2001 by Breiman and they perform with accuracy in high dimensional settings. In this article, we consider a simplified kernel-based random forest algorithm called simplified directional KeRF (Kernel…
Xinyu Chen, Dalei Yu, Xinyu Zhang
The random forest (RF) algorithm has become a very popular prediction method for its great flexibility and promising accuracy. In RF, it is conventional to put equal weights on all the base learners (trees) to aggregate their predictions. However, the predictive performances of different trees within the forest can be…
Alexander James Kilpatrick, Aleksandra Ćwiek, Shigeto Kawahara, Maki Sakamoto
'Maki Sakamoto'] This study constructs machine learning algorithms that are trained to classify samples using sound symbolism, and then it reports on an experiment designed to measure their understanding against human participants. Random forests are trained using the names of Pokémon, which are fictional video game…
Joshua Daniel Loyal, Ruoqing Zhu, Yifan Cui, Xin Zhang
Random forests are one of the most popular machine learning methods due to their accuracy and variable importance assessment. However, random forests only provide variable importance in a global sense. There is an increasing need for such assessments at a local level, motivated by applications in personalized medicine…
Tyler Kolisnik, Faeze Keshavarz-Rahaghi, Rachel Purcell, Adam Smith + 1 more
Random Forest models are widely used in genomic data analysis and can offer insights into complex biological mechanisms, particularly where features influence the target in interactive, non-linear, or non-additive ways. Currently, some of the most efficient random forest methods, in terms of computational speed, are…
Hyunwook Koh
Random Forest is a widely used tree-based ensemble learning algorithm that efficiently captures complex nonlinear relationships and higher-order feature interactions with no distributional assumptions to be satisfied. It is also well-suited to human microbiome studies, where the data are highly skewed, overdispersed…
Robert Dunne
Random Forests (RF) are a very widely used modelling tool. 34 concludes that no nonlinear model had a more widespread popularity, from health care to academia to industry, than random forests and decision trees. The bounds of the methodology are still being extended. 4 give an example with 80 million variables. It is…
Meredith L. Wallace, Lucas Mentch, Bradley J. Wheeler, Amanda L. Tapia + 5 more
'Amanda L. Tapia' 'Marc Richards' 'Siyu Zhou' 'Lixia Yi' 'Susan Redline' 'Daniel J. Buysse'] Background Machine learning tools such as random forests provide important opportunities for modeling large, complex modern data generated in medicine. Unfortunately, when it comes to understanding why machine learning models…
Zhenwei Yang, Hang Lv, Zhaofeng Xu, Xinyi Wang
Machine learning is one of the widely used techniques to pattern recognition. Use of the machine learning tools is becoming a more accessible approach for predictive model development in preventing engineering disaster. The objective of the research is to for estimation of water source using the machine learning tools.…
Moayad Alnammi, Shengchao Liu, Spencer S Ericksen, Gene E Ananiev + 6 more
Traditional small molecule drug discovery is a time consuming and costly endeavor. High-throughput chemical screening can only assess a tiny fraction of drug-like chemical space. The strong predictive power of modern machine learning methods for virtual chemical screening enables training models on known active and…
Matthew Berkowitz, Rachel MacKay Altman, Thomas M. Loughin
Few systematic comparisons of methods for constructing survival trees and forests exist in the literature. Importantly, when the goal is to predict a survival time or estimate a survival function, the optimal choice of method is unclear. We use an extensive simulation study to systematically investigate various factors…
Prashanth Athri, Vidhya Murali, Pradyumna Y Muralidhar, Cassandra Königs + 4 more
- 1. Department of Computer Science and Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham, Bengaluru, India - 2. PES Center for Pattern Recognition, Department of Computer Science and Engineering, PES University, Bengaluru, India - 3. Bioinformatics and Medical Informatics, Bielefeld University…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Authors not listed
The dual imperative of mitigating carbon emissions and maximizing hydrocarbon recovery has amplified global interest in carbon capture, utilization, and storage (CCUS) technologies. These integrated processes hold significant promise for achieving net-zero targets while extending the productive life of mature oil…
Devi Ganapathi, Wunmi Akinlemibola, Antonio Baclig, Emily Penn + 1 more
Quinones and hydroquinones are small organic molecules with numerous applications: battery electrolytes, pharmaceuticals, sensors, to name a few. An understanding of their fundamental properties, such as melting points, is essential to incorporate these compounds into relevant technologies. In this study, two different…
Authors not listed
Crystal structure prediction (CSP) is a valuable computational technique used to anticipate the likely crystal structures of a compound of interest. These methods have been proven useful in research and development of pharmaceutical solid forms and in guiding the discovery of materials with targeted properties. Despite…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…