26 papers · ranked by Valyu relevance
Baidu Li, Xinhai Li
Linear models, including t-test, ANOVA, regression, ANCOVA, and generalized linear models, are foundational tools in statistical analysis. For large datasets, such as those involving tens of thousands of genes and millions of records, numerous advanced methods have been developed to improve both computational…
Wang, Jingyuan, Ji, Jiahao
This article serves as the regression analysis lecture notes in the Intelligent Computing course cluster (including the courses of Artificial Intelligence, Data Mining, Machine Learning, and Pattern Recognition) at the School of Computer Science and Engineering, Beihang University. It aims to provide students – who are…
Ayon Roy, Tausif Al Zubayer, Nafisa Tabassum, Muhammad Nazrul Islam + 1 more
'Abdus Sattar'] Regression analysis is a well known quantitative research method that primarily explores the relationship between one or more independent variables and a dependent variable. Conducting regression analysis manually on large datasets with multiple independent variables can be tedious. An automated system…
Marijn van Vliet, Riitta Salmelin
Linear machine learning models “learn” a data transformation by being exposed to examples of input with the desired output, forming the basis for a variety of powerful techniques for analyzing neuroimaging data. However, their ability to learn the desired transformation is limited by the quality and size of the example…
Ariel I. Mundo, John R. Tipton, Timothy J. Muldoon
In biomedical research, the outcome of longitudinal studies has been traditionally analyzed using the repeated measures analysis of variance (rm-ANOVA) or more recently, linear mixed models (LMEMs). Although LMEMs are less restrictive than rm-ANOVA in terms of correlation and missing observations, both methodologies…
Bernard M. S. van Praag, J. Peter Hop, William H. Greene
In the last few decades, the study of ordinal data in which the variable of interest is not exactly observed but only known to be in a specific ordinal category has become important. To emphasize that the problem is not specific to a specific discipline we will use the neutral term coarsened observation. For…
Wilhelm Grzesiak, Daniel Zaborski, Marcin Pluciński, Magdalena Jędrzejczak-Silicka + 3 more
'Magdalena Jędrzejczak-Silicka' 'Renata Pilarczyk' 'Piotr Sablik' 'Andrea Pezzuolo'] Title: Simple Summary The current trend in animal husbandry, including cattle farming, is toward increasing stocking density and automating individual activities in animal care. Various electro-optical, acoustic, mechanical, and…
Mustafa Attallah
Pearson's correlation to select predictor variables for linear models Authors: ['Mustafa Attallah'] This article examines the limitations of Pearson's correlation in selecting predictor variables for linear models. Using mtcars and iris datasets from R, this paper demonstrates the limitation of this correlation measure…
Authors not listed
We developed OpenStats, a user-friendly web application that brings the power of the R language to researchers through a high-level interface and broad support for statistical methods such as t-tests and ANOVA. OpenStats was integrated into our electronic lab notebook Chemotion ELN via its third-party API, enabling…
Jessica I. Murphy, Nicholas E. Weaver, Audrey E. Hendricks
Longitudinal mouse models are commonly used to study possible causal factors associated with human health and disease. However, the statistical models, such as two-way ANOVA, often applied in these studies do not appropriately model the experimental design, resulting in biased and imprecise results. Here, we describe…
Noah A. Schuster, Judith J. M. Rijnhart, Lisa C. Bosman, Jos W. R. Twisk + 2 more
'Jos W. R. Twisk' 'Thomas Klausch' 'Martijn W. Heymans'] Background Confounding is a common issue in epidemiological research. Commonly used confounder-adjustment methods include multivariable regression analysis and propensity score methods. Although it is common practice to assess the linearity assumption for the…
Yu Huo, Hongpei Li, Xiao Wang, Xiaochen Du + 1 more
When analysing two-dimensional data sets, scientists are often interested in regions where one variable depends linearly on the other. Typically they use an ad hoc method to do so. Here we develop a statistically rigorous, Bayesian approach to infer the optimal partitioning of a data set into contiguous piece-wise…
Mohammad Abu-Shaira, Greg Speegle
Machine Learning requires a large amount of training data in order to build accurate models. Sometimes the data arrives over time, requiring significant storage space and recalculating the model to account for the new data. On-line learning addresses these issues by incrementally modifying the model as data is…
Aaditya Prasad Gupta
Biological systems, at all scales of organization from nucleic acids to ecosystems, are inherently complex and variable. Therefore mathematical models have become an essential tool in systems biology, linking the behavior of a system to the interaction between its components. Parameters in empirical mathematical models…
Ryan J. Murphy, Oliver J. Maclaren, Matthew J. Simpson
Throughout the life sciences, we routinely seek to interpret measurements and observations using parametrized mechanistic mathematical models. A fundamental and often overlooked choice in this approach involves relating the solution of a mathematical model with noisy and incomplete measurement data. This is often…
Andrew McCluskey
The use of mathematical transformations to reduce non-linear functions to linear problems, which can be tackled with analytical linear regression, is commonplace in the chemistry curriculum. The linearization procedure, however, assumes an incorrect statistical model for real experimental data; leading to biased…
Noah A. Schuster, Judith J. M. Rijnhart, Jos W. R. Twisk, Martijn W. Heymans
'Martijn W. Heymans'] Objective Traditional methods to deal with non-linearity in regression analysis often result in loss of information or compromised interpretability of the results. A recommended but underutilized method for modeling non-linear associations in regression models is spline functions. We explain…
Brianna C. Heggeseth, Alvaro Aleman, Hafiz T.A. Khan
There is a growing literature that suggests environmental exposure during key developmental periods could have harmful impacts on growth and development of humans. Understanding and estimating the relationship between early-life exposure and human growth is vital to studying the adverse health impacts of environmental…
Moustafa M. A. Ibrahim, Rikard Nordgren, Maria C. Kjellsson, Mats O. Karlsson
'Mats O. Karlsson'] We investigated the possible advantages of using linearization to evaluate models of residual unexplained variability (RUV) for automated model building in a similar fashion to the recently developed method “residual modeling.” Residual modeling, although fast and easy to automate, cannot identify…
Authors not listed
Nonlinear regression analysis is a popular and important tool for scientists and engineers. In this article, we introduce theories and methods of nonlinear regression and its statistical inferences using the frequentist and Bayesian statistical modeling and computation. Least squares with the Gauss-Newton method is the…
Hamdy F. F. Mahmoud
Three types of regression models researchers need to be familiar with and know the requirements of each: parametric, semiparametric and nonparametric regression models. The type of modeling used is based on how much information are available about the form of the relationship between response variable and explanatory…
Stamatia Zavitsanou, Zonghua Bo, Emanuele Casali, Matthew Langton + 1 more
Machine learning (ML) is currently transforming the field of chemistry by offering unparalleled efficiency in addressing complex challenges. Despite the progress made, a notable gap persists in the availability of user-friendly tools tailored to chemical problems involving small and sparse datasets. Here, we introduce…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European Chemicals Agency (ECHA) and US Environmental Protection Agency (EPA) have listed approximately 800k chemicals that must be further investigated for their potential environmental and/or human health risk. A significant number of these chemicals have large enough global volumes of consumption (e.g.…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European and US chemical agencies have listed approximately 800k chemicals where knowledge on potential risks to human health and the environment are lacking. Filling these data gaps experimentally is impossible so in-silico approaches and prediction are essential. Many existing models are however limited by…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European Chemicals Agency (ECHA) and US Environmental Protection Agency (EPA) have listed approximately 800k chemicals that must be further investigated for their potential environmental and/or human health risk. A significant number of these chemicals have large enough global volumes of consumption (e.g.…