22 papers · ranked by Valyu relevance
George H. John, Pat Langley
When modeling a probability distribution with a Bayesian network, we are faced with the problem of how to handle continuous variables. Most previous work has either solved the problem by discretizing, or assumed that the data are generated by a single Gaussian. In this paper we abandon the normality assumption and…
Eibe Frank, Mark Hall, Bernhard Pfahringer
Despite its simplicity, the naive Bayes classifier has surprised machine learning researchers by exhibiting good performance on a variety of learning problems. Encouraged by these results, researchers have looked to overcome naive Bayes' primary weakness�attribute independence�and improve the performance of the…
Ahmed Majid Taha, Aida Mustapha, Soong-Der Chen
When the amount of data and information is said to double in every 20 months or so, feature selection has become highly important and beneficial. Further improvements in feature selection will positively affect a wide array of applications in fields such as pattern recognition, machine learning, or signal processing.…
Sebastian Raschka
| 1 | Introduction | | 2 | | --- | --- | --- | --- | | 2 | Naive Bayes Classification | | 3 | | | 2.1 | Overview | 3 | | | 2.2 | Posterior Probabilities | 3 | | | 2.3 | Class-conditional Probabilities | 5 | | | 2.4 | Prior Probabilities | 6 | | | 2.5 | Evidence | 8 | | | 2.6 | Multinomial Naive Bayes - A Toy Example |…
Doreswamy Doreswamy, Hemanth Kolla
—In this paper, naive Bayesian and C4.5 Decision Tree Classifiers(DTC) are successively applied in materials informatics to classify the engineering materials into different classes for the selection of materials that suit the input design specifications. Here, the classifiers are analyzed individually and their…
Edith Kovács, Anna Ország, Dániel Pfeifer, András A. Benczúr
- A Generalized Naive Bayes (GNB) structure is introduced; - It is proven that the so called GNB probability distribution associated to the GNB structure gives an approximation at least as good as the probability distribution associated to the Naive Bayes structure; - Two new algorithms are introduced: Algorithm GNB-A…
Mohimenul Karim, Rashid Abid
Specific gene regions in DNA, such as cytochrome c oxidase I (COI) in animals, are defined as DNA barcodes and can be used as identifiers to distinguish species. The standard length of a DNA barcode is approximately 650 base pairs (bp). However, because of the challenges associated with sequencing technologies and the…
Napas Udomsak
—This essay investigates the question of how the naive Bayes classifier and the support vector machine compare in their ability to forecast the Stock Exchange of Thailand. The theory behind the SVM and the naive Bayes classifier is explored. The algorithms are trained using data from the month of January 2010…
Hamse Y Mussa, John BO Mitchell, Robert C Glen
Background In the last decade the standard Naive Bayes (SNB) algorithm has been widely employed in multi-class classification problems in cheminformatics. This popularity is mainly due to the fact that the algorithm is simple to implement and in many cases yields respectable classification results. Using clever…
Denice van Herwerden, Jake O'Brien, Phil Choi, Kevin Thomas + 2 more
Isotopologue identification or removal is a necessary step to reduce the number of features that need to be identified in samples analyzed with non-targeted analysis. Currently available approaches rely on either predicted isotopic patterns or an arbitrary mass tolerance, requiring information on the molecular formula…
Christopher D. Tyrrell
Simple Bayesian networks are used as classifiers for assigning an unknown something to various classes based on information about the unknown and class descriptors. Here, the case is an unknown organism being assigned a taxonomic identification. To simplify a Bayesian network, all class descriptors (characters) can be…
Sarah V Leavitt, Robyn S Lee, Paola Sebastiani, Charles R. Horsburgh + 2 more
Estimating infectious disease parameters such as the serial interval (time between symptom onset in primary and secondary cases) and reproductive number (average number of secondary cases produced by a primary case) are important to understand infectious disease dynamics. Many estimation methods require linking cases…
Pat Langley, Stephanie Sage
In this paper, we examine previous work on the naive Bayesian classifier and review its limitations, which include a sensitivity to correlated features. We respond to this problem by embedding the naive Bayesian induction scheme within an algorithm that carries out a greedy search through the space of features. We…
Concha Bielza, Pedro Larrañaga
Bayesian networks are a type of probabilistic graphical models lie at the intersection between statistics and machine learning. They have been shown to be powerful tools to encode dependence relationships among the variables of a domain under uncertainty. Thanks to their generality, Bayesian networks can accommodate…
Quang Hung Do, Jeng-Fung Chen
Classifying the student academic performance with high accuracy facilitates admission decisions and enhances educational services at educational institutions. The purpose of this paper is to present a neuro-fuzzy approach for classifying students into different groups. The neuro-fuzzy classifier used previous exam…
Asif Ali Wagan, Shahnawaz Talpur, Sanam Narejo, Wei Wang
In various fields, including medical science, datasets characterized by uncertainty are generated. Conventional clustering algorithms, designed for deterministic data, often prove inadequate when applied to uncertain data, posing significant challenges. Recent advancements have introduced clustering algorithms based on…
Zhengqiao Zhao, Alexandru Cristian, Gail Rosen
Current metagenomic taxonomic classifiers cannot computationally keep up with the pace of training data generated from genome sequencing projects, such as the exponentially-growing NCBI RefSeq bacterial genome database. When new reference sequences are added to training data, statically trained classifiers must be…
Tanvi S. Patel, Daxesh P. Patel, Mallika Sanyal, Pranav S. Shrivastav
In the present work, we examined the outcomes and accuracy of the Support vector machine (SVM) and the Naive Bayes algorithms on a dataset, to predict whether the patient has heart disease or not, and the patient’s survival prediction status. The machine learning procedures were developed using the clinically validated…
Gonzalo A. Ruz, Pamela Araya-Díaz, Pablo A. Henríquez
Background When designing a treatment in orthodontics, especially for children and teenagers, it is crucial to be aware of the changes that occur throughout facial growth because the rate and direction of growth can greatly affect the necessity of using different treatment mechanics. This paper presents a Bayesian…
A. K. M. Azad, Salem A. Alyami, Jonathan M. Keith
Bayesian networks (BNs) are widely used to model biological networks from experimental data. Many software packages exist to infer BN structures, but the chance of getting trapped in local optima is a common challenge. Some recently developed Markov Chain Monte Carlo (MCMC) samplers called the Neighborhood sampler (NS)…
Yifan Wu, Aron Walsh, Alex Ganose
What is the minimum number of experiments, or calculations, required to find an optimal solution? Relevant chemical problems range from identifying a compound with target functionality within a given phase space to controlling materials synthesis and device fabrication conditions. A common feature in this application…
Aryan Deshwal, Cory Simon, Janardhan Rao Doppa
Given a gas storage or separation task, we wish to search a library of nanoporous materials (NPMs) for the one with the optimal adsorption property. The high cost of measuring the adsorption property of an NPM, whether in the lab or a simulation, precludes exhaustive search. We explain, demonstrate, and advocate…