26 papers · ranked by Valyu relevance
Amor Keziou, Aida Toma, Nikolai Leonenko
Moment condition models are popular in statistics and econometrics, as they provide a powerful and flexible framework for estimation. However, estimation procedures based on these models can be sensitive to misspecification or the presence of outliers in the data. In the present paper, we introduce a class of robust…
Anthony Christidis, Matias Salibian-Barrera
Robust regression methods, particularly MM-estimators, are essential for analyzing datasets where heavy-tailed noise or high-leverage outliers may be present. Algorithms to compute these estimators are iterative and rely on having a good initial point. A widely-used probabilistic approach to obtaining an initial…
Peter Filzmoser, Sven Serneels, Ricardo A. Maronna, P. Van Espen
This chapter presents an introduction to robust statistics with applications of a chemometric nature. Following a description of the basic ideas and concepts behind robust statistics, including how robust estimators can be conceived, the chapter builds up to the construction (and use) of robust alternatives for some…
Aamir Raza, Mashal Talib, Muhammad Noor-ul-Amin, Nevine Gunaime + 2 more
'Imed Boukhris' 'Muhammad Nabi'] In real-life situations, we have to analyze the data that contains the atypical observations, and the presence of outliers has adverse effects on the performance of ordinary least square estimates. In this situation, redescedning M-estimators, proposed by Huber (1964), are used to…
Meng Wang, Lihua Jiang, Michael P. Snyder
With the development of high-throughput RNA sequencing (RNA-seq) technology, the Genotype Tissue-Expression (GTEx) project (4) generated a valuable resource of gene expression data from more than 11,000 samples. The large-scale data set is a powerful resource for understanding the human transcriptome. However, the…
Marco Riani, Anthony C. Atkinson, Aldo Corbellini, Domenico Perrotta
Minimum density power divergence estimation provides a general framework for robust statistics, depending on a parameter $α$, which determines the robustness properties of the method. The usual estimation method is numerical minimization of the power divergence. The paper considers the special case of linear…
Martina Sladekova, Andy P. Field, Ottavia Epifania
The general linear model (GLM) is the most frequently applied family of statistical models in psychology. Within the GLM, the effects under study are estimated using the ordinary least squares (OLS) estimation. In certain situations, OLS produces parameter estimates that are unbiased and optimal (with least possible…
Vanda M Lourenço, Joseph O Ogutu, Hans-Peter Piepho
Genomic prediction (GP) is used in animal and plant breeding to help identify the best genotypes for selection. One of the most important measures of the effectiveness and reliability of GP in plant breeding is predictive accuracy. An accurate estimate of this measure is thus central to GP. Moreover, regression models…
Harvey J Motulsky, Ronald E Brown
Background Nonlinear regression, like linear regression, assumes that the scatter of data around the ideal curve follows a Gaussian or normal distribution. This assumption leads to the familiar goal of regression: to minimize the sum of the squares of the vertical or Y-value distances between the points and the curve.…
Graciela Boente, Alejandra Mercedes Martínez
Partially linear additive models generalize linear ones since they model the relation between a response variable and covariates by assuming that some covariates have a linear relation with the response but each of the others enter through unknown univariate smooth functions. The harmful effect of outliers either in…
Aamir Raza, Muhammad Noor-ul-Amin, Amel Ayari-Akkari, Muhammad Nabi + 1 more
'Muhammad Usman Aslam'] The OLS model is built on the assumption of normality in the distribution of error terms. However, this assumption can be easily violated, especially when there are outliers in the data. A single outlier can disrupt the normality assumption of error terms, making the OLS model less effective. In…
Chun Yu, Weixin Yao, Xue Bai
Ordinary least-squares (OLS) estimators for a linear model are very sensitive to unusual values in the design space or outliers among y values. Even one single atypical value may have a large effect on the parameter estimates. This article aims to review and describe some available and popular robust techniques…
Xiaoshan Wang, Frank Emmert-Streib
The maximum likelihood estimation (MLE) method, typically used for polytomous logistic regression, is prone to bias due to both misclassification in outcome and contamination in the design matrix. Hence, robust estimators are needed. In this study, we propose such a method for nominal response data with continuous…
Elisa Cabana, Rosa E. Lillo, Henry Laniado
A robust estimator is proposed for the parameters that characterize the linear regression problem. It is based on the notion of shrinkages, often used in Finance and previously studied for outlier detection in multivariate data. A thorough simulation study is conducted to investigate: the efficiency with normal and…
Şenay Özdemir, Olçay Arslan
Ordinary least square (OLS), maximum likelihood (ML) and robust methods are the widely used methods to estimate the parameters of a linear regression model. It is well known that these methods perform well under some distributional assumptions on error terms. However, these distributional assumptions on the errors may…
Wennan Chang, Changlin Wan, Chun Yu, Weixin Yao + 2 more
Mixture regression has been widely used as a statistical model to untangle the latent subgroups of the sample population. Traditional mixture regression faces challenges when dealing with: 1) outliers and versatile regression forms; and 2) the high dimensionality of the predictors. Here, we develop an R package called…
Wesley Spiller, Neil M. Davies, Tom M. Palmer
In recent years Mendelian randomization analysis using summary data from genome-wide association studies has become a popular approach for investigating causal relationships in epidemiology. The mrrobust Stata package implements several of the recently developed methods. mrrobust is freely available as a Stata package.…
Authors not listed
Plastic mechanical recycling is the conventional technological step towards circularity. In such aspects, complex mixtures of polyolefin blends are often fed into mechanical recycling systems, resulting in moulded products with uncertain quality. To add to the difficulty of heterogeneous feedstocks, the testing of…
Peng Zheng, Ryan Barber, Reed Sorensen, Christopher Murray + 1 more
Mixed effects (ME) models inform a vast array of problems in the physical and social sciences, and are pervasive in meta-analysis. We consider ME models where the random effects component is linear. We then develop an efficient approach for a broad problem class that allows nonlinear measurements, priors, and…
Kushal K. Dey, Rahul Mazumder
Genes with correlated expression across individuals in multiple tissues are potentially informative for systemic genetic activity spanning these tissues. In this context, the tissue-level gene expression data across multiple subjects from the Genotype Tissue Expression (GTEx) Project is a valuable analytical resource.…
Van N. T. La, Stanley Nicholson, Amna Haneef, Lulu Kang + 1 more
Some data are just underappreciated. Maybe they look different or come from a different background than most other data. Maybe they don't fit neatly into common notions of what data on a ``curve'' should look like. Whatever the case, they are pigeonholed into a restricted role that limits their contributions. But if…
Authors not listed
High-throughput experimentation (HTE) in materials science generates vast, high-dimensional datasets relating synthesis parameters to material properties. While machine learning (ML) models excel at predicting properties from these parameters, they often fail to distinguish causal drivers from merely correlated…
Robert Reischke
Confidence contours in parameter space are a helpful tool to compare and classify determined estimators. For more intricate parameter estimations of non-linear nature or complex error structures, the procedure of determining confidence contours is a statistically complex task. For polymer chemists, such particular…
Esther Heid, Charles J. McGill, Florence H. Vermeire, William H. Green
Characterizing uncertainty in machine learning models has recently gained interest in the context of machine learning reliability, robustness, safety, and active learning. Here, we separate the total uncertainty into contributions from noise in the data (aleatoric) and shortcomings of the model (epistemic), further…
Eric Hermes, Khachik Sargsyan, Habib Najm, Judit Zádor
We present a new algorithm for the optimization of molecular structures to saddle points on the potential energy surface using a redundant internal coordinate system. This algorithm automates the procedure of defining the internal coordinate system, including the handling of linear bending angles, e.g. through the…
Authors not listed
Electrochemical impedance spectroscopy (EIS) coupled with distribution of relaxation times (DRT) analysis is a robust framework for characterizing electrochemical systems. However, DRT deconvolution is often plagued by spurious peaks, hindering accurate process identification and quantitative parameter estimation. To…