24 papers · ranked by Valyu relevance
Ashley Ling, El Hamidi Hay, Samuel E. Aggrey, Romdhane Rekaya + 1 more
Ordinal categorical responses are frequently collected in survey studies, human medicine, and animal and plant improvement programs, just to mention a few. Errors in this type of data are neither rare nor easy to detect. These errors tend to bias the inference, reduce the statistical power and ultimately the efficiency…
Joon Jin Song, Mohammad Arshad Rahman, Yoo-Mi Chin, James Stamey
Quantile regression extends regression analysis beyond the conditional mean, providing a richer characterization of covariate effects across the outcome distribution. For sensitive binary outcomes, however, misclassification due to underreporting can substantially bias inference. We propose a Bayesian quantile…
Jocelyn Holden Bolin, W. Holmes Finch
Statistical classification of phenomena into observed groups is very common in the social and behavioral sciences. Statistical classification methods, however, are affected by the characteristics of the data under study. Statistical classification can be further complicated by initial misclassification of the observed…
Afrah Shafquat, Ronald G. Crystal, Jason G. Mezey
Heterogeneity in definition and measurement of complex diseases in Genome-Wide Association Studies (GWAS) may lead to misdiagnoses and misclassification errors that can significantly impact discovery of disease loci. While well appreciated, almost all analyses of GWAS data consider reported disease phenotype values as…
Shannon Smith, El Hamidi Hay, Nourhene Farhat, Romdhane Rekaya
Background Misclassification has been shown to have a high prevalence in binary responses in both livestock and human populations. Leaving these errors uncorrected before analyses will have a negative impact on the overall goal of genome-wide association studies (GWAS) including reducing predictive power. A liability…
Fergus J Chadwick, Daniel T Haydon, Dirk Humseier, Jason Matthiopoulos + 1 more
Citizen science data often contain high levels of species misclassification that can bias inference and conservation decisions. Current approaches to address mislabelling rely on expert taxonomists validating every record. This approach makes intensive use of a scarce resource and reduces the role of the citizen…
Kumeren N. Govender, David W. Eyre
Culture-independent metagenomic detection of microbial species has the potential to provide rapid and precise real-time diagnostic results. However, it is potentially limited by sequencing and classification errors. We use simulated and real-world data to benchmark rates of species misclassification using 100 reference…
Michelle Xia, P. Richard Hahn
This paper considers the problem of mismeasured categorical covariates in the context of regression modeling; if unaccounted for, such misclassification is known to result in misestimation of model parameters. Here, we exploit the fact that explicitly modeling covariate misclassification leads to a mixture…
Felix Günther, Caroline Brandl, Thomas W. Winkler, Veronika Wanner + 3 more
Imaging technology and machine learning algorithms for disease classification set the stage for high-throughput phenotyping and promising new avenues for genome-wide association studies (GWAS). Despite emerging algorithms, there has been no successful application in GWAS so far. We established machine learning based…
Joseph C. Cappelleri, Richard Chambers
Introduction Quantitative patient-reported outcome (PRO) measures ideally are analyzed on their original scales and responder analyses are used to aid the interpretation of those primary analyses. As stated in the FDA PRO Guidance for Medical Product Development (2009), one way to lend meaning and interpretation to…
Leo C. McHugh, Kevin Snyder, Thomas D. Yager, Giuseppe Sartori
Uncertainty in patient classification can be measured in a number of ways, most commonly by an inter-observer agreement statistic such as Cohen’s Kappa, or by the correlation terms in a Multitrait-Multimethod Matrix. These and related statistics estimate the extent of agreement in classifying the same patients or…
Bijan Nouri, Najaf Zare, Seyyed Mohammad Taghi Ayatollahi
Background. Misclassification of exposure variables in epidemiologic studies may lead to biased estimation of parameters and loss of power in statistical inferences. In this paper, the inverse matrix method, as an efficient method of the correction of odds ratio for the misclassification of a binary exposure, was…
Shreeya Banerji
Diabetes mellitus is a growing problem, especially in developing countries. People suffering from diabetes have an increased risk of developing a number of serious health problems. Consistently high blood glucose levels can lead to serious diseases affecting the heart and blood vessels, eyes, kidney, etc. In addition…
Neal J. Thomas
Stratification in both the design and analysis of randomized clinical trials is common. Despite features in automated randomization systems to re-confirm the stratifying variables, incorrect values of these variables may be entered. These errors are often detected during subsequent data collection and verification.…
Qianhan Zeng, Yingqiu Zhu, Xuening Zhu, Feifei Wang + 4 more
'Shuning Sun' 'Meng Su' 'Hansheng Wang'] Labeling mistakes are frequently encountered in realworld applications. If not treated well, the labeling mistakes can deteriorate the classification performances of a model seriously. To address this issue, we propose an improved Na¨ıve Bayes method for text classification. It…
Lauren J. Beesley, Lars G. Fritsche, Bhramar Mukherjee
Large-scale association analyses based on observational health care databases such as electronic health records have been a topic of increasing interest in the scientific community. However, challenges of non-probability sampling and phenotype misclassification associated with the use of these data sources are often…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European Chemicals Agency (ECHA) and US Environmental Protection Agency (EPA) have listed approximately 800k chemicals that must be further investigated for their potential environmental and/or human health risk. A significant number of these chemicals have large enough global volumes of consumption (e.g.…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European Chemicals Agency (ECHA) and US Environmental Protection Agency (EPA) have listed approximately 800k chemicals that must be further investigated for their potential environmental and/or human health risk. A significant number of these chemicals have large enough global volumes of consumption (e.g.…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European and US chemical agencies have listed approximately 800k chemicals where knowledge on potential risks to human health and the environment are lacking. Filling these data gaps experimentally is impossible so in-silico approaches and prediction are essential. Many existing models are however limited by…
Ege Atacan Doğan, Peter F. Patel‐Schneider
Disjointness checks are among the most important constraint checks in a knowledge base and can be used to help detect and correct incorrect statements and internal contradictions. Wikidata is a very large, community-managed knowledge base. Because of both its size and construction, Wikidata contains many incorrect…
Authors not listed
This research delves into olfaction, a sensory modality that remains complex and inadequately understood. We aim to fill in two gaps in recent studies that attempted to use machine learning and deep learning approaches to predict human smell perception. The first one is that molecules are usually represented with…
Denice van Herwerden, Jake O'Brien, Phil Choi, Kevin Thomas + 2 more
Isotopologue identification or removal is a necessary step to reduce the number of features that need to be identified in samples analyzed with non-targeted analysis. Currently available approaches rely on either predicted isotopic patterns or an arbitrary mass tolerance, requiring information on the molecular formula…
Maria H. Rasmussen, Chenru Duan, Heather J. Kulik, Jan Halborg Jensen
With the increasingly more important role of machine learning (ML) models in chemical research, the need for putting a level of confidence to the model predictions naturally arises. Several methods for obtaining uncertainty estimates have been proposed in recent years but consensus on the evaluation of these have yet…
Sankha Subhra Mullick, Shounak Datta, Sourish Gunesh Dhekane, Swagatam Das
'Swagatam Das'] Indices quantifying the performance of classifiers under class-imbalance, often suffer from distortions depending on the constitution of the test set or the class-specific classification accuracy, creating difficulties in assessing the merit of the classifier. We identify two fundamental conditions that…