26 papers · ranked by Valyu relevance
Hélio Amante Miot
During analysis of scientific research data, it is customary to encounter anomalous values or missing data. Anomalous values can be the result of errors of recording, typing, measurement by instruments, or may be true outliers. This review discusses concepts, examples and methods for identifying and dealing with such…
Mia Hubert, Peter J. Rousseeuw, Wannes Van den Bossche
Multivariate data are typically represented by a rectangular matrix (table) in which the rows are the objects (cases) and the columns are the variables (measurements). When there are many variables one often reduces the dimension by principal component analysis (PCA), which in its basic form is not robust to outliers.…
K. K. L. B. Adikaram, M. A. Hussein, M. Effenberger, T. Becker
We introduce a new nonparametric outlier detection method for linear series, which requires no missing or removed data imputation. For an arithmetic progression (a series without outliers) with n elements, the ratio (R) of the sum of the minimum and the maximum elements and the sum of all elements is always 2/n …
Harvey J Motulsky, Ronald E Brown
Background Nonlinear regression, like linear regression, assumes that the scatter of data around the ideal curve follows a Gaussian or normal distribution. This assumption leads to the familiar goal of regression: to minimize the sum of the squares of the vertical or Y-value distances between the points and the curve.…
Eric W. Deutsch, Roger Kramer, Joseph Ames, Andrew Bauman + 21 more
Translational biomedical research is generating exponentially more data: thousands of whole-genome sequences (WGS) are now available; brain data are doubling every two years. Analyses of Big Data, including imaging, genomic, phenotypic, and clinical data, present qualitatively new challenges as well as opportunities.…
Arman Iranfar, Adriana Arza, David Atienza
— Continuous and multimodal stress detection has been performed recently through wearable devices and machine learning algorithms. However, a well-known and important challenge of working on physiological signals recorded by conventional monitoring devices is missing data due to sensors insufficient contact and…
Karim Lounici, Grégoire Pacreau
Large datasets are often affected by cell-wise outliers in the form of missing or erroneous data. However, discarding any samples containing outliers may result in a dataset that is too small to accurately estimate the covariance matrix. Moreover, the robust procedures designed to address this problem require the…
Robert M Flight, Praneeth S Bhatt, Hunter NB Moseley
Almost all correlation measures currently available are unable to directly handle missing values. Typically, missing values are either ignored completely by removing them or are imputed and used in the calculation of the correlation coefficient. In both cases, the correlation value will be impacted based on a…
D'Orazio, Marcello
This note investigates the problem of detecting outliers in longitudinal data. It compares wellknown methods used in official statistics with proposals from the fields of data mining and machine learning that are based on the distance between observations or binary partitioning trees. This is achieved by applying the…
Tamraparni Dasu, Ji Meng Loh
We introduce the notion of statistical distortion as an essential metric for measuring the effectiveness of data cleaning strategies. We use this metric to propose a widely applicable yet scalable experimental framework for evaluating data cleaning strategies along three dimensions: glitch improvement, statistical…
Vanda M Lourenço, Joseph O Ogutu, Hans-Peter Piepho
Genomic prediction (GP) is used in animal and plant breeding to help identify the best genotypes for selection. One of the most important measures of the effectiveness and reliability of GP in plant breeding is predictive accuracy. An accurate estimate of this measure is thus central to GP. Moreover, regression models…
Martin Seifrid, Stanley Lo, Dylan Choi, Gary Tom + 12 more
Martin Seifrid 1 , Stanley Lo 2 , Dylan G. Choi 3 , Gary Tom 2 , My Linh Le 3 , Kunyu Li 3 , Rahul Sankar 3 , Hoai-Thanh Vuong 3 , Hiba Wakidi 3 , Ahra Yi 3 , Ziyue Zhu 3 , Nora Schopp 3 , Aaron Peng 3 , Benjamin Luginbuhl 3 , Thuc-Quyen Nguyen 3 , Alán Aspuru-Guzik 2
Pete R. Jones
This paper considers how best to identify statistical outliers in psychophysical datasets, where the underlying sampling distributions are unknown. Eight methods are described, and each is evaluated using Monte Carlo simulations of a typical psychophysical experiment. The best method is shown to be one based on a…
Yannik Schälte, Emad Alamoudi, Jan Hasenauer
Approximate Bayesian Computation (ABC) is a likelihood-free parameter inference method for complex stochastic models in systems biology and other research areas. While conceptually simple, its practical performance relies on the ability to efficiently compare relevant features in simulated and observed data via…
Matthew S. Smith, Karthik Devarajan
A problem that arises frequently in high-throughput biological studies is the assessment of technical reproducibility of data obtained under homogeneous experimental conditions. This is an important problem considering the significant growth in the number of high-throughput technologies that have become available to…
W. Holmes Finch
The presence of outliers can very problematic in data analysis, leading statisticians to develop a wide variety of methods for identifying them in both the univariate and multivariate contexts. In case of the latter, perhaps the most popular approach has been Mahalanobis distance, where large values suggest an…
Claudio Agostinelli, Andy Leung, Vı́ctor J. Yohai, Ruben H. Zamar
of cellwise and casewise contamination Authors: ['Claudio Agostinelli' 'Andy Leung' 'Vı́ctor J. Yohai' 'Ruben H. Zamar'] Multivariate location and scatter matrix estimation is a cornerstone in multivariate data analysis. We consider this problem when the data may contain independent cellwise and casewise outliers. Flat…
Zeynel Cebeci, Cagatay Cebeci, Yalcin Tahtali, Lutfi Bayyurt + 1 more
'Yilun Shang'] Outliers are data points that significantly deviate from other data points in a data set because of different mechanisms or unusual processes. Outlier detection is one of the intensively studied research topics for identification of novelties, frauds, anomalies, deviations or exceptions in addition to…
Authors not listed
Accurate prediction of melting points for pure molecules remains a significant challenge in predictive chemistry, with implications across various scientific fields, including materials science, drug discovery, and separations chemistry. Traditional methods, such as group contribution (GC) techniques, have shown…
Marjorie Fonnesu, Nicola Kuczewski
While the utilisation of different methods of outliers correction has been shown to counteract the inferential error produced by the presence of contaminating data not belonging to the studied population; the effects produced by their utilisation when samples do not contain contaminating outliers are less clear. Here a…
Freedom N Gumedze, Dan Jackson
Background Meta-analysis typically involves combining the estimates from independent studies in order to estimate a parameter of interest across a population of studies. However, outliers often occur even under the random effects model. The presence of such outliers could substantially alter the conclusions in a…
Ghayath Janoudi, Mara Uzun (Rada), Deshayne B. Fell, Joel G. Ray + 5 more
'Angel M. Foster' 'Randy Giffen' 'Tammy Clifford' 'Mark C. Walker' 'Raymond Francis Sarmiento'] Clinical discoveries largely depend on dedicated clinicians and scientists to identify and pursue unique and unusual clinical encounters with patients and communicate these through case reports and case series. This process…
Authors not listed
Predicting solution conformation and aggregation of conjugated polymers remains a bottleneck for translating solution processing into controlled film microstructure and for closing the loop in self-driving laboratories. We construct a cleaned, machine-readable dataset of 256 entries that links polymer size…
Authors not listed
The discovery of chemically novel or structurally anomalous metal-organic frameworks (MOFs) is essential for expanding reticular design space and enhancing dataset reliability. We present CHEM-AD (Chemically Unusual Metal–organic Frameworks via Autoencoder-based Detection), a label-free, CPU-efficient pipeline that…
Diba Behnoudfar, Cory Simon, Joshua Schrier
Aqueous, two-phase systems (ATPSs) may form upon mixing two solutions of independently water-soluble compounds. Many separation, purification, and extraction processes rely on ATPSs. Predicting the miscibility of solutions can accelerate and reduce the cost of the discovery of new ATPSs for these applications. Whereas…
Authors not listed
Plastic mechanical recycling is the conventional technological step towards circularity. In such aspects, complex mixtures of polyolefin blends are often fed into mechanical recycling systems, resulting in moulded products with uncertain quality. To add to the difficulty of heterogeneous feedstocks, the testing of…