15 papers · ranked by Valyu relevance
Joon Jin Song, Mohammad Arshad Rahman, Yoo-Mi Chin, James Stamey
Quantile regression extends regression analysis beyond the conditional mean, providing a richer characterization of covariate effects across the outcome distribution. For sensitive binary outcomes, however, misclassification due to underreporting can substantially bias inference. We propose a Bayesian quantile…
Augustine Denteh, Désiré Kédagni
The difference-in-differences (DID) design is one of the most popular methods used in empirical economics research. However, there is almost no work examining what the DID method identifies in the presence of a misclassified treatment variable. This paper studies the identification of treatment effects in DID designs…
Michelle Xia, P. Richard Hahn
This paper considers the problem of mismeasured categorical covariates in the context of regression modeling; if unaccounted for, such misclassification is known to result in misestimation of model parameters. Here, we exploit the fact that explicitly modeling covariate misclassification leads to a mixture…
Christopher Meek
Understanding prediction errors and determining how to fix them is critical to building effective predictive systems. In this paper, we delineate four types of prediction errors (mislabeling, representation, learner and boundary errors) and demonstrate that these four types characterize all prediction errors. In…
Nathan TeBlunthuis, Valerie Hase, Chung‐hong Chan
Automated classifiers (ACs), often built via supervised machine learning (SML), can categorize large, statistically powerful samples of data ranging from text to images and video. They have become widely popular measurement devices in communication science and related fields. Despite this popularity, even highly…
Kimberly A. Hochstedler, Martin T. Wells
In biomedical and public health association studies, binary outcome variables may be subject to misclassification, resulting in substantial bias in effect estimates. The feasibility of addressing binary outcome misclassification in regression models is often hindered by model identifiability issues. In this paper, we…
Neal J. Thomas
Stratification in both the design and analysis of randomized clinical trials is common. Despite features in automated randomization systems to re-confirm the stratifying variables, incorrect values of these variables may be entered. These errors are often detected during subsequent data collection and verification.…
Qianhan Zeng, Yingqiu Zhu, Xuening Zhu, Feifei Wang + 4 more
'Shuning Sun' 'Meng Su' 'Hansheng Wang'] Labeling mistakes are frequently encountered in realworld applications. If not treated well, the labeling mistakes can deteriorate the classification performances of a model seriously. To address this issue, we propose an improved Na¨ıve Bayes method for text classification. It…
Michael R. Smith, Tony Martinez
Removing or filtering outliers and mislabeled instances prior to training a learning algorithm has been shown to increase classification accuracy. A popular approach for handling outliers and mislabeled instances is to remove any instance that is misclassified by a learning algorithm. However, an examination of which…
Augustine Denteh, Pierre E. Nguimkeu
This paper considers the estimation of binary choice models when survey responses are possibly misclassified but one of the response category can be validated. Partial validation may occur when survey questions about participation include follow-up questions on that particular response category. In this case, we show…
Quinten Meertens, Cees Diks, H.J. van den Herik, Frank W. Takes
National statistical institutes currently investigate how to improve the output quality of official statistics based on machine learning algorithms. A key obstacle is concept drift, i.e., when the joint distribution of independent variables and a dependent (categorical) variable changes over time. Under concept drift…
Hiroyuki Kasahara, Katsumi Shimotsu
We study identification in nonparametric regression models with a misclassified and endogenous binary regressor when an instrument is correlated with misclassification error. We show that the regression function is nonparametrically identified if one binary instrument variable and one binary covariate satisfy the…
Ege Atacan Doğan, Peter F. Patel‐Schneider
Disjointness checks are among the most important constraint checks in a knowledge base and can be used to help detect and correct incorrect statements and internal contradictions. Wikidata is a very large, community-managed knowledge base. Because of both its size and construction, Wikidata contains many incorrect…
Tristan Mary‐Huard, Vittorio Perduca, Gilles Blanchard, Marie‐Laure Martin‐Magniette
'Marie‐Laure Martin‐Magniette'] In the context of finite mixture models one considers the problem of classifying as many observations as possible in the classes of interest while controlling the classification error rate in these same classes. Similar to what is done in the framework of statistical test theory…
Sankha Subhra Mullick, Shounak Datta, Sourish Gunesh Dhekane, Swagatam Das
'Swagatam Das'] Indices quantifying the performance of classifiers under class-imbalance, often suffer from distortions depending on the constitution of the test set or the class-specific classification accuracy, creating difficulties in assessing the merit of the classifier. We identify two fundamental conditions that…