22 papers · ranked by Valyu relevance
Guoji Fu, Chengbin Hou, Xin Yao
—The topological information is essential for studying the relationship between nodes in a network. Recently, Network Representation Learning (NRL), which projects a network into a low-dimensional vector space, has been shown their advantages in analyzing large-scale networks. However, most existing NRL methods are…
Rishikesh U. Kulkarni, Catherine L. Wang, Carolyn R. Bertozzi
We report Hierarch, a Python package to perform hypothesis tests and compute confidence intervals on hierarchical experimental designs. Using a combination of permutation resampling and bootstrap aggregation, Hierarch can be used to perform hypothesis tests that maintain nominal Type I error rates and generate…
Philippe Aubry
Title: Graphical abstract
Rishikesh U. Kulkarni, Catherine L. Wang, Carolyn R. Bertozzi, Dina Schneidman-Duhovny
'Dina Schneidman-Duhovny'] While hierarchical experimental designs are near-ubiquitous in neuroscience and biomedical research, researchers often do not take the structure of their datasets into account while performing statistical hypothesis tests. Resampling-based methods are a flexible strategy for performing these…
Udo Boehm, Maarten Marsman, Dora Matzke, Eric-Jan Wagenmakers
Psychological experiments often yield data that are hierarchically structured. A number of popular shortcut strategies in cognitive modeling do not properly accommodate this structure and can result in biased conclusions. To gauge the severity of these biases, we conducted a simulation study for a two-group experiment.…
Christian Damgaard
In many applied cases of ecological and environmental modelling, there is a sizeable variation among the measured variables due to measurement- and sampling error. Such measurement- and sampling error among the independent variables may lead to regression dilution and biased prediction intervals in traditional…
Nathan Kirk, Ivan Gvozdanović, Sonja Petrović
This paper proposes a multilevel sampling algorithm for fiber sampling problems in algebraic statistics, inspired by Henry Wynn's suggestion to adapt multilevel Monte Carlo (MLMC) ideas to discrete models. Focusing on log-linear models, we sample from high-dimensional lattice fibers defined by algebraic constraints.…
Sebastian Krumscheid, Per Pettersson
Quantifyingthe effect of uncertainties in computationally complex systems where only point evaluations in the stochastic domain but no regularity conditions are available is limited to sampling-based techniques. This work presents an adaptive sequential stratification estimation method that uses Latin Hypercube…
Martin Robinson, Alan Bond, Alexandr Simonov, Jie Zhang + 1 more
Recently, we have introduced the use of techniques drawn from Bayesian statistics to recover kinetic and thermodynamic parameters from voltammetric data, and were able to show that the technique of large amplitude ac voltammetry yielded significantly more accurate parameter values than the equivalent dc approach. In…
Michael D. Shields, Jiaxin Zhang
Latin hypercube sampling (LHS) is generalized in terms of a spectrum of stratified sampling (SS) designs referred to as partially stratified sample (PSS) designs. True SS and LHS are shown to represent the extremes of the PSS spectrum. The variance of PSS estimates is derived along with some asymptotic properties. PSS…
H. R. N. van Erp, R. O. Linger, Pieter van Gelder
Stated differently, the θ are not directly observable, they can only be inferred. So, if we wish to assign a function u to θ, then we have to take our uncertainty, in regards to the actual value of the vector θ, into account. But if we do so, then this will give us highly dimensional and highly intractable integrals.…
Samuel R. Lucas
The multilevel model has become a staple of social research. I textually and formally explicate sample design features that, I contend, are required for unbiased estimation of macro-level multilevel model parameters and the use of tools for statistical inference, such as standard errors. After detailing the limited and…
Augustijn A.A. de Boer, Seyed Mostafa Kia, Saige Rutherford, Mariam Zabihi + 7 more
Normative modelling is an emerging technique for parsing heterogeneity in clinical cohorts. This can be implemented in practice using hierarchical Bayesian regression, which provides an elegant probabilistic solution to handle site variation in a federated learning framework. However, applications of this method to…
Authors not listed
Accurate and efficient calculation of alchemical free energies is a critical challenge in computational chemistry, frequently hindered by the inherent limitations of conventional Thermodynamic Integration (TI) methods. These limitations include poor phasespace overlap between discrete alchemical states, inefficient…
Alec Wong, Angela Fuller, J. Andrew Royle
Rare species present challenges to data collection, particularly when the species is spatially clustered over large areas, such that the encounter frequency of the organism is low. Sampling where the organism is absent consumes resources, and offers relatively low-quality information which are often difficult to model…
June Gorostidi, Adem Ait, Jordi Cabot, Javier Luis Izquierdo
Software repositories is one of the sources of data in Empirical Software Engineering, primarily in the Mining Software Repositories field, aimed at extracting knowledge from the dynamics and practice of software projects. With the emergence of social coding platforms such as GitHub, researchers have now access to…
Kanwal Iqbal, Syed Muhammad Muslim Raza, Tahir Mahmood, Muhammad Riaz + 1 more
'Muhammad Riaz' 'Mohamed R. Abonazel'] Advancements in sensor technology have brought a revolution in data generation. Therefore, the study variable and several linearly related auxiliary variables are recorded due to cost-effectiveness and ease of recording. These auxiliary variables are commonly observed as…
Emmanuel Ren, François-Xavier Coudert
Molecular adsorption in nanoporous materials has many large-scale industrial applications ranging from separation to storage. To design the best materials, computational simulations are key in guiding the experimentation and engineering processes. Because nanoporous materials exist in a plethora of forms, we need to…
Michel H Hof, Anita CJ Ravelli, Aeilko H Zwinderman
Background In population-based observational studies, non-participation and delayed response to the invitation to participate are complications that often arise during the recruitment of a sample. When both are not properly dealt with, the composition of the sample can be different from the desired composition.…
Zenabu Suboi, Thomas J. Hladish, Wim Delva, C. Marijn Hazelbag
Complex models are often fitted to data using simulation-based calibration, a computationally challenging process. Several calibration methods to improve computational efficiency have been developed with no consensus on which methods perform best. We did a simulation study comparing the performance of 5 methods that…
Yao Xiao, Kang Fu, Kun Li
The sampling importance resampling method is widely utilized in various fields, such as numerical integration and statistical simulation. In this paper, two modified methods are presented by incorporating two variance reduction techniques commonly used in Monte Carlo simulation, namely antithetic sampling and Latin…
Zhimian Hao, Chonghuan Zhang, Alexei Lapkin
We propose a workflow for reduction in the time required for data generation during generation of statistical digital twins. This methodology is particularly relevant for real-world engineering problems when data generation is expensive. A prerequisite for building surrogates is sufficient input/output data, whereas…