25 papers · ranked by Valyu relevance
Bernhard Schölkopf, Julius von Kügelgen
We describe basic ideas underlying research to build and understand artificially intelligent systems: from symbolic approaches via statistical learning to interventional models relying on concepts of causality. Some of the hard open problems of machine learning and AI are intrinsically related to causality, and…
Tatsuya Daikoku, Kevin Kamermans, Maiko Minatoya
Statistical learning starts at an early age and is intimately linked to brain development and the emergence of individuality. Through such a long period of statistical learning, the brain updates and constructs statistical models, with the model’s individuality changing based on the type and degree of stimulation…
Ernest Fokoué
The rapid ascent of artificial intelligence (AI) is often portrayed as a revolution born from computer science and engineering. This narrative, however, obscures a fundamental truth: the theoretical and methodological core of AI is, and has always been, statistical. This paper systematically argues that the field of…
Aryan Yazdanpanah, Michael Chong Wang, Ethan Trepka, Marissa Benz + 1 more
Natural environments are abundant with patterns and regularities. These regularities can be captured through statistical learning, which strongly influences perception, memory, and other cognitive functions. By combining a sequence-prediction task with an orthogonal multidimensional reward learning task, we tested…
Sumio Watanabe
between Statistical Mechanics and Machine Learning Theory Authors: ['Sumio Watanabe'] Mathematical equivalence between statistical mechanics and machine learning theory has been known since the 20th century, and researches based on such equivalence have provided novel methodology in both theoretical physics and…
Orsolya Pesthy, Zsuzsanna Viktória Pesthy, Teodóra Vékony, Karolina Janacsek + 2 more
The brain must balance the automatic extraction of environmental regularities with top-down cognitive control, yet the causal neural mechanisms governing this interplay are debated. In particular, the hemispheric contributions of the dorsolateral prefrontal cortex (DLPFC) remain unresolved. Here, we applied inhibitory…
Joseph Andersen
Statistical models have seen a significant rise in popularity in recent years. Despite their undeniable success in various industry use cases such as sabermetrics, investment portfolio management, and artificial intelligence, there has been immense debate about the value of results produced by statistical methods. This…
Samuel K. Sheppard, Nicolas Arning, David W. Eyre, Daniel J. Wilson
The availability of large genome datasets has changed the microbiology research landscape. Analyzing such data requires computationally demanding analyses, and new approaches have come from different data analysis philosophies. Machine learning and statistical inference have overlapping knowledge discovery aims and…
Gili Lior, Yuval Shalev, Gabriel Stanovsky, Ariel Goldstein
The human brain is an adaptive learning system that can generalize to new tasks and unfamiliar environments. The traditional view is that such adaptive behavior requires a structural change of the learning system (e.g., via neural plasticity). In this work, we use artificial neural networks, specifically large language…
Giovanni Cerulli
We present two related Stata modules, r ml stata and c ml stata, for fitting popular Machine Learning (ML) methods both in a regression and a classification setting. Using the recent Stata/Python integration platform (sfi) of Stata 16, these commands provide hyper-parameters' optimal tuning via K-fold cross-validation…
Joram Soch, Carsten Allefeld
We propose the statistical modelling approach to supervised learning (i.e. predicting labels from features) as an alternative to algorithmic machine learning (ML). The approach is demonstrated by employing a multivariate general linear model (MGLM) describing the effects of labels on features, possibly accounting for…
Juan Jovel, Russell Greiner
Machine learning (ML) approaches are a collection of algorithms that attempt to extract patterns from data and to associate such patterns with discrete classes of samples in the data-e.g., given a series of features describing persons, a ML model predicts whether a person is diseased or healthy, or given features of…
Marie Salditt, Theresa Eckes, Steffen Nestler
Psychotherapy has been proven to be effective on average, though patients respond very differently to treatment. Understanding which characteristics are associated with treatment effect heterogeneity can help to customize therapy to the individual patient. In this tutorial, we describe different meta-learners, which…
Yiru Jiang, Jing Luo, Danqing Huang, Ya Liu + 1 more
Microorganisms play an important role in natural material and elemental cycles. Many common and general biology research techniques rely on microorganisms. Machine learning has been gradually integrated with multiple fields of study. Machine learning, including deep learning, aims to use mathematical insights to…
Fernando Marmolejo‐Ramos, Mauricio Tejo, Marek Brabec, Jakub Kuzilek + 8 more
The advent of technological developments is allowing to gather large amounts of data in several research fields. Learning analytics (LA)/educational data mining has access to big observational unstructured data captured from educational settings and relies mostly on unsupervised machine learning (ML) algorithms to make…
Diego Marcondes, Ulisses Braga-Neto
We propose generalized resubstitution error estimators for regression, a broad family of estimators, each corresponding to a choice of empirical probability measures and loss function. The usual sum of squares criterion is a special case corresponding to the standard empirical probability measure and the quadratic…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Narjice Chafai, Ichrak Hayah, Isidore Houaga, Bouabid Badaoui
The advent of modern genotyping technologies has revolutionized genomic selection in animal breeding. Large marker datasets have shown several drawbacks for traditional genomic prediction methods in terms of flexibility, accuracy, and computational power. Recently, the application of machine learning models in animal…
Authors not listed
We developed OpenStats, a user-friendly web application that brings the power of the R language to researchers through a high-level interface and broad support for statistical methods such as t-tests and ANOVA. OpenStats was integrated into our electronic lab notebook Chemotion ELN via its third-party API, enabling…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
S M Mehedi Zaman, Wasay Mahmood Qureshi, Md. Mohsin Sarker Raihan, Abdullah Bin Shams + 1 more
'Abdullah Bin Shams' 'Sharmin Sultana'] Abstract— cardiovascular disease, especially heart failure is one of the major health hazard issues of our time and is a leading cause of death worldwide. Advancement in data mining techniques using machine learning (ML) models is paving promising prediction approaches. Data…
Zongben Xu, Jun Shu, Deyu Meng
This paper introduces a ‘simulating learning methodology’ (SLeM) approach for the learning methodology determination in general and for Auto6 ML in particular, and reports the SLeM framework, approaches, algorithms and applications.
Authors not listed
Quantitative Structure Activity Relationship (QSAR) remains an effective tool for early-stage chemical modelling and virtual screening in drug design. The advancements in this field are led by two core paradigms, 1) descriptor engineering, where complex fixed-length vectors of compounds are generated and conventional…
Authors not listed
Quantitative Structure-Activity Relationship (QSAR) modeling is a pillar of computational drug discovery. However, standard machine learning (ML) models are often confounded by the high-dimensional and intensely correlated nature of molecular descriptors. A model may identify a "bulk" property (e.g., molecular weight)…
Ricardo Stefani
The use of data science, artificial intelligence, and big data in the field of chemistry has recently grown to speed up the discovery of new materials, drugs, and synthetic substances and the identification of automated compounds. Machine learning and data science are commonly used in organic chemistry to predict…