21 papers · ranked by Valyu relevance
Joram Soch, Carsten Allefeld
We propose the statistical modelling approach to supervised learning (i.e. predicting labels from features) as an alternative to algorithmic machine learning (ML). The approach is demonstrated by employing a multivariate general linear model (MGLM) describing the effects of labels on features, possibly accounting for…
Lineu Alberto Cavazani de Freitas, Wagner Hugo Bonat
This article describes the R package htmcglm implemented for performing hypothesis tests on regression and dispersion parameters of multivariate covariance generalized linear models (McGLMs). McGLMs provide a general statistical modeling framework for normal and non-normal multivariate data analysis along with a wide…
Omid Chatrabgoun, Alireza Daneshkhah, Parisa Torkaman, Mark Johnston + 3 more
'Nader Sohrabi Safa' 'Ali Kashif Bashir' 'Zakariya Yahya Algamal'] Many machine learning techniques have been used to construct gene regulatory networks (GRNs) through precision matrix that considers conditional independence among genes, and finally produces sparse version of GRNs. This construction can be improved…
Geoffrey R. Hosack
A statistical method for the elicitation of priors in Bayesian generalised linear models (GLMs) and extensions is proposed. Probabilistic predictions are elicited from the expert to parametrise a multivariate t prior distribution for the unknown linear coefficients of the GLM and an inverse gamma prior for the…
Fentaw Abegaz, Davar Abedini, Lemeng Dong, Johan A. Westerhuis + 3 more
In microbiome studies, addressing the unique characteristics of sequence data-such as compositionality, zero inflation, overdispersion, high dimensionality, and non-normality-is crucial for accurate analysis. In addition, integrating experimental design elements into microbiome data analysis is important for…
M. Gomtsyan, Céline Lévy-Leduc, Sarah Ouadah, Laure Sansonnet + 2 more
'Christophe Bailly' 'Loïc Rajjou'] Abstract. We propose a novel and efficient iterative two-stage variable selection approach for multivariate sparse GLARMA models, which can be used for modelling multivariate discrete-valued time series. Our approach consists in iteratively combining two steps: the estimation of the…
Shao-Hsuan Wang, Ray Bai, Hsin‐Hsiung Huang
In recent years, the literature on Bayesian high-dimensional variable selection has rapidly grown. It is increasingly important to understand whether these Bayesian methods can consistently estimate the model parameters. To this end, shrinkage priors are useful for identifying relevant signals in high-dimensional data.…
Anita Brobbey, Samuel Wiebe, Alberto Nettel-Aguirre, Colin Bruce Josephson + 3 more
generalized estimation equations Authors: ['Anita Brobbey' 'Samuel Wiebe' 'Alberto Nettel-Aguirre' 'Colin Bruce Josephson' 'Tyler Williamson' 'Lisa M Lix' 'Tolulope T. Sajobi'] Discriminant analysis procedures that assume parsimonious covariance and/or means structures have been proposed for distinguishing between two…
Dan Wang, Jun Teng, Changheng Zhao, Xinhao Zhang + 5 more
Current methods of multivariate analysis require complete multivariate phenotypes from each individual and have a computational time complexity of O(n^2^) per SNP, where n is the sample size. We develop an efficient genomic multivariate analysis tool (GMAT) for genome-wide association studies of multiple correlated…
Guilherme Parreira da Silva, Henrique Aparecido Laureano, Ricardo Rasmussen Petterle, Paulo Justiniano Ribeiro + 1 more
'Ricardo Rasmussen Petterle' 'Paulo Justiniano Ribeiro' 'Wagner Hugo Bonat'] Researchers are often interested in understanding the relationship between a set of covariates and a set of response variables. To achieve this goal, the use of regression analysis, either linear or generalized linear models, is largely…
Fabio Morgante, Peter Carbonetto, Gao Wang, Yuxin Zou + 2 more
Predicting phenotypes from genotypes is a fundamental task in quantitative genetics. With technological advances, it is now possible to measure multiple phenotypes in large samples. Multiple phenotypes can share their genetic component; therefore, modeling these phenotypes jointly may improve prediction accuracy by…
Tien-Wen Lee
The General Linear Model (GLM) has been widely used in research, where error term has been treated as noise. However, compelling evidence suggests that in biological systems, the target variables may possess their innate variances. A modified GLM was proposed to explicitly model biological variance and non-biological…
Rezzy Eko Caraka, Rung-Ching Chen, Su-Wen Huang, Shyue-Yow Chiou + 2 more
'Prana Ugiana Gio' 'Bens Pardamean'] Background In heart data mining and machine learning, dimension reduction is needed to remove multicollinearity. Meanwhile, it has been proven to improve the interpretation of the parameter model. In addition, dimension reduction can also increase the time of computing in high…
Maeve McGillycuddy, Gordana Popović, Benjamin M. Bolker, David I. Warton
'David I. Warton'] Multivariate random effects with unstructured variance-covariance matrices of large dimensions, q, can be a major challenge to estimate. In this paper, we introduce a new implementation of a reduced-rank approach to fit large dimensional multivariate random effects by writing them as a linear…
Chin-Sheng Teng, Xuesong Wang, Cheng Liu, Qishan Wang + 2 more
Genome-wide association studies (GWAS) often analyze one trait at a time, but multivariate GWAS can increase the power by leveraging trait correlations. However, existing methods struggle with high computational demands, especially when analyzing over five traits. We present EMmvGWAS, an efficient multivariate GWAS…
Xynthia Kavelaars, Joris Mulder, Maurits Kaptein
Background In medical, social, and behavioral research we often encounter datasets with a multilevel structure and multiple correlated dependent variables. These data are frequently collected from a study population that distinguishes several subpopulations with different (i.e., heterogeneous) effects of an…
Sanjeena Subedi, Utkarsh J. Dang
Modern biological data are often multivariate discrete counts, and there has been a dearth of statistical distributions to directly model such counts in an efficient manner. While mixed Poisson distributions, e.g., negative binomial distribution, are often the distribution of choice for univariate data, multivariate…
Yuanqing Lu, Timur Fazletdinov, Zhiwen Pan, Katrin Wondraczek + 1 more
The synthesis of nanoscale particles and particle aggregates from liquid or gaseous precursors is affected by a variety of trade-off relations, for example, in terms of product composition, yield, or energy efficiency. Machine-supported process evaluation and learning (ML) of these relations enables optimization…
Authors not listed
We developed OpenStats, a user-friendly web application that brings the power of the R language to researchers through a high-level interface and broad support for statistical methods such as t-tests and ANOVA. OpenStats was integrated into our electronic lab notebook Chemotion ELN via its third-party API, enabling…
Authors not listed
Inverse problems, where we seek the values of inputs to a model that lead to a desired set of outputs, are a challenges subset of problems in science and engineering. In this work we demonstrate the use of two generative AI methods to solve inverse problems. We compare this approach to two more conventional approaches…
Authors not listed
Ensuring the trustworthiness of machine learning (ML) models in high-stake applications is crucial. One such application is predicting anti-cancer drug sensitivity, where ML models are built with the final goal of integrating them into treatment recommendation systems for personalized medicine. Here, we propose a…