26 papers · ranked by Valyu relevance
Michael J. Crosse, Nathaniel J. Zuk, Giovanni M. Di Liberto, Aaron R. Nidiffer + 2 more
'Aaron R. Nidiffer' 'Sophie Molholm' 'Edmund C. Lalor'] Cognitive neuroscience, in particular research on speech and language, has seen an increase in the use of linear modeling techniques for studying the processing of natural, environmental stimuli. The availability of such computational tools has prompted similar…
Hristos Tyralis, Georgia Papacharalampous
Although the Kling-Gupta efficiency ($\mathrm{KGE}$) is widely adopted for model evaluation in hydrology, its properties as a statistical estimator remain unexplored. Investigating these properties is necessary because parameter estimation and forecast evaluation are inherently linked. To address this, we formalize the…
Wang, Jingyuan, Ji, Jiahao
This article serves as the regression analysis lecture notes in the Intelligent Computing course cluster (including the courses of Artificial Intelligence, Data Mining, Machine Learning, and Pattern Recognition) at the School of Computer Science and Engineering, Beihang University. It aims to provide students – who are…
Ayon Roy, Tausif Al Zubayer, Nafisa Tabassum, Muhammad Nazrul Islam + 1 more
'Abdus Sattar'] Regression analysis is a well known quantitative research method that primarily explores the relationship between one or more independent variables and a dependent variable. Conducting regression analysis manually on large datasets with multiple independent variables can be tedious. An automated system…
Bernard M. S. van Praag, J. Peter Hop, William H. Greene
In the last few decades, the study of ordinal data in which the variable of interest is not exactly observed but only known to be in a specific ordinal category has become important. To emphasize that the problem is not specific to a specific discipline we will use the neutral term coarsened observation. For…
Wilhelm Grzesiak, Daniel Zaborski, Marcin Pluciński, Magdalena Jędrzejczak-Silicka + 3 more
'Magdalena Jędrzejczak-Silicka' 'Renata Pilarczyk' 'Piotr Sablik' 'Andrea Pezzuolo'] Title: Simple Summary The current trend in animal husbandry, including cattle farming, is toward increasing stocking density and automating individual activities in animal care. Various electro-optical, acoustic, mechanical, and…
Mustafa Attallah
Pearson's correlation to select predictor variables for linear models Authors: ['Mustafa Attallah'] This article examines the limitations of Pearson's correlation in selecting predictor variables for linear models. Using mtcars and iris datasets from R, this paper demonstrates the limitation of this correlation measure…
Authors not listed
We developed OpenStats, a user-friendly web application that brings the power of the R language to researchers through a high-level interface and broad support for statistical methods such as t-tests and ANOVA. OpenStats was integrated into our electronic lab notebook Chemotion ELN via its third-party API, enabling…
Lincoln Huber, Boris Kovalerchuk, Charles Recaido
Understanding black-box Machine Learning methods on multidimensional data is a key challenge in Machine Learning. While many powerful Machine Learning methods already exist, these methods are often unexplainable or perform poorly on complex data. This paper proposes visual knowledge discovery approaches based on…
Ethan M. McCormick, Michelle L. Byrne, John C. Flournoy, Kathryn L. Mills + 1 more
Longitudinal data are becoming increasingly available in developmental neuroimaging. To maximize the promise of this wealth of information on how biology, behavior, and cognition change over time, there is a need to incorporate broad and rigorous training in longitudinal methods into the repertoire of developmental…
Yu Huo, Hongpei Li, Xiao Wang, Xiaochen Du + 1 more
When analysing two-dimensional data sets, scientists are often interested in regions where one variable depends linearly on the other. Typically they use an ad hoc method to do so. Here we develop a statistically rigorous, Bayesian approach to infer the optimal partitioning of a data set into contiguous piece-wise…
Mohammad Abu-Shaira, Greg Speegle
Machine Learning requires a large amount of training data in order to build accurate models. Sometimes the data arrives over time, requiring significant storage space and recalculating the model to account for the new data. On-line learning addresses these issues by incrementally modifying the model as data is…
Andrew McCluskey
The use of mathematical transformations to reduce non-linear functions to linear problems, which can be tackled with analytical linear regression, is commonplace in the chemistry curriculum. The linearization procedure, however, assumes an incorrect statistical model for real experimental data; leading to biased…
Noah A. Schuster, Judith J. M. Rijnhart, Jos W. R. Twisk, Martijn W. Heymans
'Martijn W. Heymans'] Objective Traditional methods to deal with non-linearity in regression analysis often result in loss of information or compromised interpretability of the results. A recommended but underutilized method for modeling non-linear associations in regression models is spline functions. We explain…
Lance F. Merrick, Dennis N. Lozada, Xianming Chen, Arron H. Carter
Most genomic prediction models are linear regression models that assume continuous and normally distributed phenotypes, but responses to diseases such as stripe rust (caused by Puccinia striiformis f. sp. tritici) are commonly recorded in ordinal scales and percentages. Disease severity (SEV) and infection type (IT)…
David R. Heit, Waldemar Ortiz‐Calo, Mairi K. P. Poisson, Andrew R. Butler + 1 more
'Andrew R. Butler' 'Remington J. Moll'] Title: Abstract Generalized linear models (GLMs) are an integral tool in ecology. Like general linear models, GLMs assume linearity, which entails a linear relationship between independent and dependent variables. However, because this assumption acts on the link rather than the…
Nicholas C Chesnaye, Merel van Diepen, Friedo Dekker, Carmine Zoccali + 2 more
'Carmine Zoccali' 'Kitty J Jager' 'Vianda S Stel'] Title: ABSTRACT True linear relationships are rare in clinical data. Despite this, linearity is often assumed during analyses, leading to potentially biased estimates and inaccurate conclusions. In this introductory paper, we aim to first describe-in a non-mathematical…
Luis F. Arias-Giraldo, Marlon E. Cobos
Here, we present the new R package “enmpa,” which includes a range of tools for modeling ecological niches using presence-absence data via logistic generalized linear models. The package allows users to calibrate, select, project, and evaluate models using independent data. We have emphasized a comprehensive search for…
Aaditya Prasad Gupta
Biological systems, at all scales of organization from nucleic acids to ecosystems, are inherently complex and variable. Therefore mathematical models have become an essential tool in systems biology, linking the behavior of a system to the interaction between its components. Parameters in empirical mathematical models…
Rex Parsons, Oliver Jayasinghe, Nicole White, Prasad Chunduri + 1 more
The complexity, volume, and importance of time series data across various research domains highlight the necessity for tools that can efficiently analyze, visualize, and extract insights. Cosinor modeling is a widely used methodology to estimate or compare rhythmic characteristics in time series datasets. Time series…
Satwik Acharyya, Debdeep Pati, Dipankar Bandyopadhyay, Shumei Sun
Beta distributions are commonly used to model proportion valued response variables, commonly encountered in longitudinal studies. In this article, we develop semi-parametric Beta regression models for proportion valued responses, where the aggregate covariate effect is summarized and flexibly modeled, using a…
Stamatia Zavitsanou, Zonghua Bo, Emanuele Casali, Matthew Langton + 1 more
Machine learning (ML) is currently transforming the field of chemistry by offering unparalleled efficiency in addressing complex challenges. Despite the progress made, a notable gap persists in the availability of user-friendly tools tailored to chemical problems involving small and sparse datasets. Here, we introduce…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European and US chemical agencies have listed approximately 800k chemicals where knowledge on potential risks to human health and the environment are lacking. Filling these data gaps experimentally is impossible so in-silico approaches and prediction are essential. Many existing models are however limited by…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European Chemicals Agency (ECHA) and US Environmental Protection Agency (EPA) have listed approximately 800k chemicals that must be further investigated for their potential environmental and/or human health risk. A significant number of these chemicals have large enough global volumes of consumption (e.g.…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Saer Samanipour, Jake O'Brien, Malcolm Reid, Kevin Thomas + 1 more
The European Chemicals Agency (ECHA) and US Environmental Protection Agency (EPA) have listed approximately 800k chemicals that must be further investigated for their potential environmental and/or human health risk. A significant number of these chemicals have large enough global volumes of consumption (e.g.…