Search · four archives
Search · four archives
19 papers · ranked by Valyu relevance
Davide Chicco, Giuseppe Jurman
Background To evaluate binary classifications and their confusion matrices, scientific researchers can employ several statistical rates, accordingly to the goal of the experiment they are investigating. Despite being a crucial issue in machine learning, no widespread consensus has been reached on a unified elective…
Yuki Itaya, Jun Tamura, Kenichi Hayashi, Kouji Yamamoto
Evaluating classifications is crucial in statistics and machine learning, as it influences decision-making across various fields, such as patient prognosis and therapy in critical conditions. The Matthews correlation coefficient (MCC), also known as the phi coefficient, is recognized as a performance metric with high…
Yuki Itaya, Junsuke Tamura, Kenichi Hayashi, Kouji Yamamoto
Evaluating classifications is crucial in statistics and machine learning, as it influences decisionmaking across various fields, such as patient prognosis and therapy in critical conditions. The Matthews correlation coefficient (MCC) is recognized as a performance metric with high reliability, offering a balanced…
Davide Chicco, Giuseppe Jurman
A binary classification is a computational procedure that labels data elements as members of one or another category. In machine learning and computational statistics, input data elements which are part of two classes are usually encoded as 0’s or -1’s (negatives) and 1’s (positives). During a binary classification, a…
Jun Tamura, Yuki Itaya, Kenichi Hayashi, Kouji Yamamoto
Classification problems are essential statistical tasks that form the foundation of decision-making across various fields, including patient prognosis and treatment strategies for critical conditions. Consequently, evaluating the performance of classification models is of significant importance, and numerous evaluation…
Milton Pividori, Marylyn D. Ritchie, Diego H. Milone, Casey S. Greene
Correlation coefficients are widely used to identify patterns in data that may be of particular interest. In transcriptomics, genes with correlated expression often share functions or are part of disease-relevant biological processes. Here we introduce the Clustermatch Correlation Coefficient (CCC), an efficient…
Hélio Amante Miot
Transformation of data (for example, logarithmic, square root, 1/x) in order to obtain a normal distribution to enable Pearson’s coefficient to be tested is a valid option for samples with asymmetrical data distributions ( [gf0200] : V1 x V6). However, it should be borne in mind that, in common with techniques that…
Benjamín M. Taylor
Pearson's correlation is an important summary measure of the amount of dependence between two variables. It is natural to want to generalise the concept of correlation as a single number that measures the inter-relatedness of three or more variables e.g. how 'correlated' are a collection of variables in which non are…
M. Baak, R. F. Koopman, H. L. Snoek, S. Klous
A prescription is presented for a new and practical correlation coefficient, φK, based on several refinements to Pearson's hypothesis test of independence of two variables. The combined features of φK form an advantage over existing coefficients. First, it works consistently between categorical, ordinal and interval…
Jianji Wang, Nanning Zheng
Multivariate correlation analysis plays an important role in various fields such as statistics, economics, and big data analytics. In this paper, we propose a pair of measures, multivariate correlation coefficient (MCC) and multivariate uncorrelation coefficient (MUC), to measure the strength of the correlation and…
Zhihao Yao, Jing Zhang, Xiufen Zou
Background With the advance of high throughput sequencing, high-dimensional data are generated. Detecting dependence/correlation between these datasets is becoming one of most important issues in multi-dimensional data integration and co-expression network construction. RNA-sequencing data is widely used to construct…
E. Coissac, C. Gonindard-Melodelima
Molecular biology and ecology studies can produce high dimension data. Estimating correlations and shared variation between such data sets are an important step in disentangling the relationships between different elements of a biological system. Unfortunately, classical approaches are susceptible to producing falsely…
Edoardo Saccenti, Margriet H. W. B. Hendriks, Age K. Smilde
Correlation coefficients are abundantly used in the life sciences. Their use can be limited to simple exploratory analysis or to construct association networks for visualization but they are also basic ingredients for sophisticated multivariate data analysis methods. It is therefore important to have reliable estimates…
Ben O’Neill
In this review article we consider linear regression analysis from a geometric perspective, looking at standard methods and outputs in terms of the lengths of the relevant vectors and the angles between these vectors. We show that standard regression output can be written in terms of the lengths and angles between the…
Bin Zhuo, Duo Jiang, Yanming Di
When a statistical test is repeatedly applied to rows of a data matrix—such as in differential-expression analysis of gene expression data, correlations among data rows will give rise to correlations among corresponding test statistic values. Correlations among test statistic values create many inferential challenges…
Sébastien Buczinski, Paul Mills
Simple Summary Veterinary science is based on data collection at the animal or herd level. Beyond the variability in the variable in question, the data collected can depend on the device used or the person performing the measurement. Determination of these sources of variation is crucial to be able to use these…
Liming Zhao, Huting Wang, yingsheng zhang
In recent years, machine learning (ML) models have been found to quickly predict various molecular properties with accuracy comparable to high level quantum chemistry methods. One such example is the calculation of electrostatic potential (ESP). Different ESP prediction ML models were proposed to generate surface…
Hao Tang, Tianle Yue, Ying Li
Machine learning (ML) has become an important technique in materials science, markedly accelerating the discovery and design of novel materials, and concurrently lowering the burden of experimental costs. Uncertainty quantification (UQ) plays a pivotal role in the accurate prediction and innovative design of novel…
Mohamad Mohebifar, Erin R. Johnson, Christopher Rowley
The exchange-hole dipole moment (XDM) model from density-functional theory predicts atomic and molecular London dispersion coefficients from first principles, providing an innovative strategy to validate the dispersion terms of molecular-mechanical force fields. In this work, the XDM model was used to obtain the London…