23 papers · ranked by Valyu relevance
Wang, Jingyuan, Ji, Jiahao
This article serves as the regression analysis lecture notes in the Intelligent Computing course cluster (including the courses of Artificial Intelligence, Data Mining, Machine Learning, and Pattern Recognition) at the School of Computer Science and Engineering, Beihang University. It aims to provide students – who are…
Gowtham Rajendiran, Jebakumar Rethnaraj, Shrikant Zade, Ramakrishna Guttula + 1 more
Aeroponic vertical tower farming is a cost-effective, sustainable method for optimizing the food crop-Lactuca Sativa (lettuce-a greeny leaf vegetable); yet accurate biomass prediction of the lettuce crop remains challenging due to the non-linear relationship between the climatic conditions and the variable lettuce…
Muni Lakshmi G K, Mokesh Rayalu G
Air pollution, especially elevated particulate matter concentrations, presents a substantial risk to public health and environmental sustainability in urban regions. By employing machine learning and hybrid ensemble models, this study develops a robust frame work for predicting the Air Quality Index (AQI). A multi-step…
Alexander Hsu, Zhaiming Shen, Wenjing Liao, Rongjie Lai
Pre-trained transformers are able to learn from examples provided as part of the prompt without any weight updates, a remarkable ability known as in-context learning (ICL). Despite its demonstrated efficacy across various domains, the theoretical understanding of ICL is still developing. Whereas most existing theory…
Tianjian Qin, Koen van Benthem, Luis Valente, Rampal Etienne
Reconstructing the forces that shaped macroevolutionary histories from extant phylogenies is fundamentally challenging: richly parameterized diversification models are often only weakly identifiable; different evolutionary mechanisms can yield nearly indistinguishable tree shapes. Here we use a model with evolutionary…
Theodoros Anagnostopoulos, Evanthia Zervoudi, Christos Anagnostopoulos, Apostolos Christopoulos + 1 more
Linear regression analysis focuses on predicting a numeric regressand value based on certain regressor values. In this context, k-Nearest Neighbors (k-NN) is a common non-parametric regression algorithm, which achieves efficient performance when compared with other algorithms in literature. In this research effort an…
Cheyenne N. Jarman, Taal Levi, Mark Novak
Applications of machine learning in ecology are rapidly expanding. Symbolic regression is gaining particular attention for its success in reverse-engineering human-readable explanatory population models, including the logistic growth and Lotka-Volterra equations, from simulated and laboratory-based population time…
Francesco Freni, A. Fries, Linus Kühne, Markus Reichstein + 1 more
We consider a regression setting where observations are collected in different environments modeled by different data distributions. The field of out-of-distribution (OOD) generalization aims to design methods that generalize better to test environments whose distributions differ from those observed during training.…
Authors not listed
Supervised deep learning has become a standard approach to deliver competitive predictive tools that allow relating the structure of molecules and their physicochemical features to properties such as binding to protein targets, performance as electronic materials, and reactivity. However, efforts to understand how…
Authors not listed
Bonkowski and De Souza [Sol. Stat. Ionics 429, 116967 (2025)] provide a guide for performing molecular dynamics simulations of ion transport, including methods for estimating diffusion coefficients and their uncertainties from mean-squared displacement (MSD) data. The discussion of uncertainty in estimated diffusion…
Theodoros Anagnostopoulos, Evanthia K. Zervoudi, Christos Anagnostopoulos, Apostolos Christopoulos + 1 more
Linear regression analysis focuses on predicting a numeric regressand value based on certain regressor values. In this context, k-Nearest Neighbors (k-NN) is a common non-parametric regression algorithm, which achieves efficient performance when compared with other algorithms in literature. In this research effort an…
Authors not listed
Phase equilibrium calculations are crucial in chemical engineering design and optimization processes. The PC-SAFT equation of state (EoS) can precisely calculate phase equilibrium, but is relatively complex and computationally intensive. Surrogate models are mathematically simple models that map or regress the…
Authors not listed
Metal hydrides play a pivotal role in a wide range of applications, including hydrogen storage, compression, heat management, and catalysis, making them a central focus of interdisciplinary research spanning chemistry, materials science, and engineering. The performance of the metal hydride based systems is strongly…
Mathias Bourel
Decision trees are one of the fundamental tools in statistical learning due to their interpretability, flexibility, and their ability to adapt to nonlinear structures. Among them, the Classification and Regression Trees, introduced by Breiman, Friedman, Olshen, and Stone in 1984, became one of the most influential…
Fabian Woller, Paul Martini, Souptik Sen, David B. Blumenthal + 1 more
Gene regulatory networks (GRNs) are graph-based representations of regulatory relationships between transcription factors and target genes. Various tools exist to infer GRNs from gene expression data, but since this task is computationally intensive, statistical significance estimates are often omitted. While…
William Casey, Leigh Metcalf, Shirshendu Chatterjee, Heeralal Janwa + 3 more
Many real-world problems feature nonlinear dynamic processes. Classical mathematical models may be adequate to describe a single dynamic process in isolation, but can be easily undermined by two natural and simple kinds of phenomenological variations: the emergence (or activation) of an additional dynamic process, and…
Authors not listed
Terminally labeled DNA oligonucleotides have wide applications in modern biology and biotechnological applications. It has been observed that the fluorescent intensity of light released from these fluorescent labels is heavily influenced by the terminal sequence of nucleotides. Recent studies have assayed and published…
Francesco G. Rinaldi, Eugenio Piasini
To make sense of a noisy world, living beings constantly face decisions between competing interpretations for ambiguous sensory data. This process parallels statistical model selection, where most frameworks, like the Akaike Information Criterion (AIC) and the Bayesian Information Criterion (BIC), are based on a…
Antony Mizzi, David M. Walker, Michael Small, José F. F. Mendes
We derive a penalty strength criterion for ridge regression using stochastic complexity, which is a refined variant of the minimum description length principle. Since stochastic complexity does not typically account for the effect of regularization on complexity, despite its ability to simplify models, we are required…
Edossa Merga Terefe, Merga Abdissa Aga
Vehicle insurance claim severity modeling requires accurate and interpretable methods that can handle skewed and heterogeneous loss data. This study provides a structured empirical comparison between classical parametric regression models and tree-based ensemble learning approaches for predicting claim size conditional…
Allen Bush-Beaupré, Simon Coroller-Chouraki, Marc Bélisle
Much ecological research focuses on phenomena where a given variable can affect another either directly or indirectly through the effect of one or more intervening variables. While various methods to quantify the magnitude of these effects are available, they can be difficult to interpret in a meaningful way…
Authors not listed
The rapid growth of worldwide computing power has transformed in silico chemistry into a discipline that is integrated into the daily work of many chemists. Nowadays, researchers find it increasingly straightforward to predict a wide range of molecular properties and chemi- cal processes at reasonable computational…
Jinwoo Lee, Junghoon Justin Park, Maria Pak, Seung Yun Choi + 1 more
Analyzing individual differences in treatment or exposure effects is a central challenge in psychology and behavioral sciences. Conventional statistical models have focused on average treatment effects, overlooking individual variability, and struggling to identify key moderators. Generalized Random Forest (GRF) can…