23 papers · ranked by Valyu relevance
Fentaw Abegaz, Davar Abedini, Lemeng Dong, Johan A. Westerhuis + 3 more
In microbiome studies, addressing the unique characteristics of sequence data-such as compositionality, zero inflation, overdispersion, high dimensionality, and non-normality-is crucial for accurate analysis. In addition, integrating experimental design elements into microbiome data analysis is important for…
Will Penny, Tom Sambrook, Louis Renoult
Factorial designs are a mainstay of the scientific paradigm, allowing the effects of multiple experimental factors and their interactions to be efficiently studied within a single experiment. In brain imaging, however, multivariate data analyses commonly proceed using multivariate decoding and we argue that the…
Pekka Korhonen, Francis K.C. Hui, Jenni Niku, Sara Taskinen + 2 more
Background Over the past decade, joint species distribution models (JSDMs) and model-based ordination have emerged as powerful tools for the analysis of community ecology data. Generalized linear latent variable models (GLLVMs) offer a flexible framework for multivariate analysis of a wide range of data types, based on…
Yusrianti Hanike, Purhadi, Achmad Choiruddin
Regression modeling for multivariate count data often struggles with assumption of overdispersion and correlation among response variables. To address these issues, this study proposes a new model called Multivariate Correlated Poisson Generalized Inverse Gaussian Regression (MCPGIGR), which integrates random effects…
Maksym Hrachov, Hans-Peter Piepho, Niaz Md. Farhat Rahman, Waqas Ahmed Malik
Key message Several seemingly distinct regression methods are closely related. Environmental covariates delivered improved prediction, and a new approach improves estimation of prediction variance. Abstract In plant breeding and variety testing, there is an increasing interest in making use of environmental information…
Giovanni Toto, Peter Müller, Abhra Sarkar
We propose a flexible Bayesian approach for estimating the joint density of a multivariate outcome of interest in the presence of categorical covariates. Leveraging a Gaussian copula framework, our method effectively captures the dependence structure across different coordinates of the multivariate response. The…
Angela Andreella, Livio Finos
Linear mixed models are widely used to analyze non-independent data, but inference for fixed effects can be unreliable under misspecification of the random-effects distribution, inaccurate Fisher information estimation, or convergence failures, leading to a lack of control over false positives. These difficulties are…
Javier de la Fuente, Mijke Rhemtulla, Travis T. Mallard, Michel Nivard + 2 more
Many medical, physiological, and psychiatric traits and disorders are highly polygenic and exhibit complex patterns of genetic sharing and differentiation. In 2018, we introduced Genomic Structural Equation Modelling (Genomic SEM) as a formal framework and free, open source, R-based software for modelling the…
Sarah S. Ji, Benjamin B. Chu, Hua Zhou, Kenneth Lange + 1 more
Copulas, generalized estimating equations, and generalized linear mixed models promote the analysis of grouped data where non-normal responses are correlated. Unfortunately, parameter estimation remains challenging in these three frameworks. Based on prior work of Tonda, we derive a new class of probability density…
Benjamin Christoffersen, Keith Humphreys, Alessandro Gasparini, Birzhan Akynkozhayev + 2 more
Joint models are well suited to modelling linked data from laboratories and health registers. However, there are few examples of joint models that allow for (a) multiple markers, (b) multiple survival outcomes (including terminal events, competing events, and recurrent events), (c) delayed entry and (d) scalability. We…
Veronica Vinciotti, Ernst C. Wit
Contingency tables are the canonical representation of multivariate categorical data. As the size of the contingency table grows exponentially with the number of variables, even a moderate number of variables, each with a moderate number of levels, results in a huge number of cells, the majority of which remains empty…
Johanna Wilroth, Nancy Sotero Silva, Ali Tafakkor, Bruno de Avo Mesquita + 4 more
Functional near infrared spectroscopy (fNIRS) is increasingly used in hearing and communication research, with advantages such as robustness to movement artifacts, improved spatial resolution, and flexibility of contexts in which it can be applied. At the same time, the field is progressively moving towards more…
Sören Budig, Charlotte Vogel, Frank Schaarschmidt
Overdispersion, a common issue in clustered multinomial data, can lead to biased standard errors and compromised statistical inference if not adequately addressed. This study describes a comprehensive procedure for constructing multiple comparisons of interest and applying multiplicity adjustments in the analysis of…
Jacinta Guirguis, Luke E.B. Goodyear, Daniel Pincheira-Donoso
Phylogenetic modelling has consolidated as the analytical standard to address hypotheses about the patterns and dynamics of biodiversity in inter-specific contexts. These analyses are traditionally performed implementing phylogenetic linear models where single outcomes are regressed against multiple predictors without…
Anna Ly, Rune Haubo Bojesen Christensen, Douglas Bates, Martin Maechler + 1 more
The lme4 R package can be used to fit generalized linear mixed models (GLMMs), which extend the class of linear mixed models (LMMs). The two main extensions provided by GLMMs are (1) allowing for the conditional distribution of the response given the random effects to be non-Gaussian (e.g. binomial, Poisson) and (2)…
Razvan G Romanescu, Michelle Liu
We consider the problem of optimal testing for genetic interaction between two variants, allowing for possible main effects. Finding a most powerful test is important because it ends a series of attempts in the literature to construct ever more powerful tests for interaction at the variant pair level. Testing under a…
Coralie Williams, Maeve McGillycuddy, Szymon M. Drobniak, Benjamin M. Bolker + 2 more
Phylogenetic generalised linear mixed models (PGLMMs) help ecologists to distinguish ecological drivers from other processes shaping evolutionary patterns, yet existing implementations are often limited in distributional scope or computational speed. We compare five R packages for fitting PGLMMs and highlight the new…
Xiang Li, Mirko Signorelli
Generalized linear mixed models (GLMMs) are widely used for analyzing correlated data, such as longitudinal and multilevel data. With over 15 $\texttt{R}$ packages available on $\texttt{CRAN}$ for fitting GLMMs, practitioners face a difficult choice regarding which package yields accurate estimates, converges reliably…
Xuelei Wang, Jana Zweerings, Michael Lührs, Fengyu Cong + 5 more
Identifying informative voxels is a critical, yet challenging step in functional magnetic resonance imaging (fMRI), particularly for multivariate analyses involving multiple related conditions. Existing approaches often rely on predefined regions of interest (ROIs) or activation-based criteria, which may be…
Authors not listed
Quantitative Structure-Activity Relationship (QSAR) modeling is a pillar of computational drug discovery. However, standard machine learning (ML) models are often confounded by the high-dimensional and intensely correlated nature of molecular descriptors. A model may identify a "bulk" property (e.g., molecular weight)…
Authors not listed
Lipidomics provides critical insights into disease mechanisms, biomarker discovery, and precision medicine, but conventional LC-MS workflows often require large sample volumes, limiting their application in minimally invasive studies. Here, we establish a nanoflow liquid chromatography (nanoLC) platform coupled with…
Authors not listed
Decades of extensive research have proved that β-amyloid (Aβ) peptides and their aggregation, inducing oxidative stress in the brain, play a key role in Alzheimer’s disease (AD) development. Moreover, Aβ peptides bind to Cu(II) ions, and the resulting complexes accelerate the aggregation process while promoting the…
Authors not listed
High-throughput experimentation (HTE) in materials science generates vast, high-dimensional datasets relating synthesis parameters to material properties. While machine learning (ML) models excel at predicting properties from these parameters, they often fail to distinguish causal drivers from merely correlated…