23 papers · ranked by Valyu relevance
Mona Azadkia, Sourav Chatterjee
We propose a coefficient of conditional dependence between two random variables Y and Z given a set of other variables X1, . . . , Xp, based on an i.i.d. sample. The coefficient has a long list of desirable properties, the most important of which is that under absolutely no distributional assumptions, it converges to a…
Gordana C. Popovic, David I. Warton, Fiona J. Thomson, Francis K. C. Hui + 1 more
Ecologists often investigate co-occurrence patterns in multi-species data in order to gain insight into the ecological causes of observed co-occurrences. Apart from direct associations between two species, two species may co-occur because they both respond in similar ways to environmental variables, or due to the…
Bilol Banerjee
This article deals with the problem of testing conditional independence between two random vectors X and Y given a confounding random vector Z. Several authors have considered this problem for multivariate data. However, most of the existing tests has poor performance against local contiguous alternatives beyond linear…
Joanne K Daggy, Huiping Xu, Siu L Hui, Roland E Gamache + 1 more
'Shaun J Grannis'] Background Methods for linking real-world healthcare data often use a latent class model, where the latent, or unknown, class is the true match status of candidate record-pairs. This commonly used model assumes that agreement patterns among multiple fields within a latent class are independent. When…
Maria Bolsinova, Dylan Molenaar
The most common process variable available for analysis due to tests presented in a computerized form is response time. Psychometric models have been developed for joint modeling of response accuracy and response time in which response time is an additional source of information about ability and about the underlying…
Dries Debeer, Carolin Strobl
Background Random forest based variable importance measures have become popular tools for assessing the contributions of the predictor variables in a fitted random forest. In this article we reconsider a frequently used variable importance measure, the Conditional Permutation Importance (CPI). We argue and illustrate…
Lei Zan, Anouar Meynaoui, Charles K. Assaad, Emilie Devijver + 2 more
'Eric Gaussier' 'Antonio M. Scarfone'] In this study, we focus on mixed data which are either observations of univariate random variables which can be quantitative or qualitative, or observations of multivariate random variables such that each variable can include both quantitative and qualitative components. We first…
Anne-Laure Boulesteix, Silke Janitza, Alexander Hapfelmeier, Kristel Van Steen + 1 more
'Kristel Van Steen' 'Carolin Strobl'] In an interesting and quite exhaustive review on Random Forests (RF) methodology in bioinformatics Touw et al. address-among other topics-the problem of the detection of interactions between variables based on RF methodology. We feel that some important statistical concepts, such…
Alex E. Yuan, Wenying Shou
In disciplines from biology to climate science, a routine task is to compute a correlation between a pair of time series, and determine whether the correlation is statistically significant (i.e. unlikely under the null hypothesis that the time series are independent). This problem is challenging because time series…
Suzanne H. Keddie, Oliver Baerenbold, Ruth H. Keogh, John Bradley
Background Latent class models are increasingly used to estimate the sensitivity and specificity of diagnostic tests in the absence of a gold standard, and are commonly fitted using Bayesian methods. These models allow us to account for ‘conditional dependence’ between two or more diagnostic tests, meaning that the…
P-A. G. Maugis
Entries of datasets are often collected only if an event occurred: taking a survey, enrolling in an experiment and so forth. However, such partial samples bias classical correlation estimators. Here we show how to correct for such sampling effects through two complementary estimators of event conditional correlation…
Alex E Yuan, Wenying Shou
Complex ecosystems are challenging to understand as they often defy manipulative experiments for practical or ethical reasons. In response, several fields have developed parallel approaches to infer causal relations from observational time series. Yet these methods are easy to misunderstand and often controversial.…
Hu Jun, Xianggui Qu
In this article we provide a substantial discussion on the statistical concept of conditional independence, which is not routinely mentioned in most elementary statistics and mathematical statistics textbooks. Under the assumption of conditional independence, an extended version of Bayes' Theorem is then proposed with…
Emilio Salinas, Terrence R. Stanford, Nicholas V Swindale
Intuitively, combining multiple sources of evidence should lead to more accurate decisions than considering single sources of evidence individually. In practice, however, the proper computation may be difficult, or may require additional data that are inaccessible. Here, based on the concept of conditional…
Luis M. de Campos, Serafı́n Moral
In this paper we study different concepts of independence for convex sets of probabilities. There will be two basic ideas for independence. The first is irrelevance. Two variables are independent when a change on the knowledge about one variable does not affect the other. The second one is factorization. Two variables…
Florian Griessenberger, Wolfgang Trutschnig, Robert R. Junker
Correlations belong to the standard repertoire of ecologists for quantifying the strength of dependence between two random variables. Classical dependence measures are usually not capable of detecting non-monotonic or non-functional dependencies. Furthermore, they completely fail to detect asymmetry and direction in…
Wilhelmiina Hämäläinen, Geoffrey I. Webb
Statistically sound pattern discovery harnesses the rigour of statistical hypothesis testing to overcome many of the issues that have hampered standard data mining approaches to pattern discovery. Most importantly, application of appropriate statistical tests allows precise control over the risk of false discoveries –…
Necla Koçhan, Gözde Yazgı Tütüncü, Göknur Giner
Recent developments in the next-generation sequencing (NGS) based on RNA-sequencing (RNA-Seq) allow researchers to measure the expression levels of thousands of genes for multiple samples simultaneously. In order to analyze these kind of data sets, many classification models have been proposed in the literature. Most…
Avi Pfeffer
Suppose we are given the conditional probability of one variable given some other variables. Normally the full joint distribution over the conditioning variables is required to determine the probability of the conditioned variable. Under what circumstances are the marginal distributions over the conditioning variables…
Luke Fenton-Glynn
Joseph Halpern and Judea Pearl ([23]) draw upon structural equation models to develop an attractive analysis of ‘actual cause’. Their analysis is designed for the case of deterministic causation. I show that their account can be naturally extended to provide an elegant treatment of probabilistic causation. 1. 1…
Sara Giarrusso, Paola Gori-Giorgi, Federica Agostini
We generalize the definitions of local scalar potentials named vkin and vN−1, which are relevant to properly describe phenomena such as molecular dissociation with density-functional theory, to the case in which the electronic wavefunction corresponds to a complex current-carrying state. In such a case, an extra term…
Authors not listed
Two kinetic schemes of the general modifier mechanism have been analysed in a quasi-steady state approximation, assuming that the reaction product concentration is negligible (a natural assumption for the initial rate method) and without additional simplifying assumptions. The characteristic equations have been…
Xinqiang Ding, John Drohan
A common approach for computing free energy differences among multiple states is to build a perturbation graph connecting the states and compute free energy differences on all edges of the graph. Such perturbation graphs are often designed to have cycles. Because free energy is a function of states, the free energy…