25 papers · ranked by Valyu relevance
Rong Huang, Su Tao
Automated Machine Learning (AutoML) aims to streamline the end-to-end process of ML models, yet current approaches remain constrained by rigid rule-based frameworks and structured input requirements that create barriers for non-expert users. Despite advances in Large Language Models (LLMs) demonstrating capabilities in…
Syed Mehedi Hasan Nirob, Shamim Ehsan, Moqsadur Rahman, Summit Haque
—Large language models (LLMs) have made it remarkably easy to synthesize plausible source code from natural language prompts. While this accelerates software development and supports learning, it also raises new risks for academic integrity, authorship attribution, and responsible AI use. This paper investigates the…
Melih Peker, Ozcan Ozturk
Selecting a good set of optimization flags requires extensive effort and expert input. While most of the prior research considers using static, spatial, or dynamic features, some of the latest research directly applied deep neural networks to source code. We combined the static features, spatial features, and deep…
Matteo De Matola, Giorgio Arcara
Convolutional neural networks (CNNs) are a class of artificial neural networks (ANNs). Since the early 2010s, they have been widely adopted as models of primate vision and classifiers of neuroimaging data, becoming relevant for a wealth of neuroscientific fields. However, the majority of neuroscience researchers come…
Martin Prause
Despite the growing popularity of AI coding assistants, over 80% of machine learning (ML) projects fail to deliver real business value. This study creates and tests a Machine Learning Canvas, a practical framework that combines business strategy, software engineering, and data science in order to determine the factors…
M. Shahbaz Ismail, Sara Shahzad, Fahmi H. Quradaa, Sajid Anwar
Semantic code clone detection plays an essential role in software maintenance and quality assurance, as it helps uncover fragments of code that express the same logic even when their syntax has been altered or deliberately obfuscated. In this study, we propose a framework that combines hybrid representation learning…
Vlastimil Martinek, Andrea Gariboldi, Dimosthenis Tzimotoudis, Mark Galea + 7 more
Extracting knowledge from biomedical data is crucial for advancing our understanding of biological systems and developing novel therapeutics. The quantity, quality, and resolution of biomedical data constantly evolves, requiring the automation of biomedical machine learning (ML). Existing Automated ML tools lack…
Meryem Gharmili, Youssef Aatif, Alj Abdelkamel
Generative AI coding assistants are increasingly used to write machine-learning code, yet their ability to produce reliable LSTM implementations for financial prediction remains underexplored. This study evaluates the LSTM code generated by seven assistants ChatGPT 4.5, GitHub Copilot, Deepseek 3, Perplexity, Gemini…
Huifang Ma, Zhicheng Ji, Tara Al-Hashimy, Austin Allen + 31 more
Title: Significance Large language models (LLMs) are increasingly used in science and engineering, yet their real-world effectiveness in data analysis remains unclear. In this study, graduate students used LLMs to tackle biomedical data challenges on Kaggle, a popular data science platform. Despite limited programming…
Paulo Lyra, Junhao Qiu, Khai Dang, Alyssa Pybus + 6 more
Machine learning is increasingly central to biomedical research, but using machine learning well often requires substantial computational expertise and methodological care to produce high-quality results. To make machine learning tools more accessible to biomedical researchers while supporting best-practice approaches…
Vlastimil Martinek, Andrea Gariboldi, Dimosthenis Tzimotoudis, Mark Galea + 7 more
The past decades have witnessed the transformation of molecular biology into a truly data-driven science, in large part due to the growth in the quantity and variety of molecular biology data generated by technologies such as mass spectrometry and high-throughput sequencing (, ). An important step in the analysis…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Benedikt Fein, Gordon Fraser
The trend of embedding source code for machine learning applications also enables new opportunities in learning analytics in programming education, but which code embedding approach is most suitable for learning analytics remains an open question. A common approach to embedding source code lies in treating the code as…
Eric W. Bridgeford, Iain Declan Campbell, Zijiao Chen, Zhicheng Lin + 4 more
While AI coding tools have demonstrated potential to accelerate software development, their use in scientific computing raises critical questions about code quality and scientific validity. In this paper, we provide twelve practical tips for AI-assisted coding that balance the capabilities of AI with the demands of…
Elitsa Yotkova, Violeta Kastreva, Dimitar Dimitrov, Ivan Koychev + 1 more
SemEval-2026 Task 13 investigates machine-generated code detection across multiple programming languages and application scenarios, asking participating systems to generalize to unseen languages and domains. This paper describes our participation in Subtask A (binary classification) and explores both pretrained code…
Hung Q. Vo, Huy Q. Vo, Son T. Ly, Zhihao Wan + 5 more
Conventional tissue image analysis software provides foundational capabilities for cellular analysis, including segmentation, basic morphological feature extraction, and spatial organization analysis. However, these tools often require manual intervention and are not well integrated with code-driven automation…
Alexandros Vassiliades, Nikolaos Polatidis, Stamatios Samaras, Sotiris Diplaris + 4 more
This study explores the explainability capabilities of large language models (LLMs), when employed to autonomously generate machine learning (ML) solutions. We examine two classification tasks: (i) a binary classification problem focused on predicting driver alertness states, and (ii) a multilabel classification…
David S. Fischer
Agentic AI is increasingly deployed on complex problems, often using chain-of-thought prompting to ground predictions in stepwise reasoning. In biomedical research, assistive agents could make this reasoning accessible to human scientists: for example, intermediate conclusions could be critically evaluated based on…
Gabriel Poesia, Georgia Gabriela Sampaio
Classical models for supervised machine learning, such as decision trees, are efficient and interpretable predictors, but their quality is highly dependent on the particular choice of input features. Although neural networks can learn useful representations directly from raw data (e.g., images or text), this comes at…
Authors not listed
Recent advances in machine learning force fields (MLFF) have significantly extended the reach of atomistic simulations. Continuous progress in this field requires reliable reference datasets, accurate MLFF architectures, and efficient active learning strategies to enable robust modeling of complex molecular and…
Authors not listed
Realizing the promise of artificial intelligence (AI) to accelerate scientific progress and deliver technological impact depends on how effectively AI can be integrated into real-world decision- making processes. As Peter Norvig states, “Somewhat remarkably, almost all AI research until very recently has assumed that…
Authors not listed
LC-HRMS is widely used in forensic toxicology for broad-scope screening. When a newly emerging or rarely encountered compound is tentatively identified, toxicologists must decide whether it may be relevant to the case and, if so, quantify it. Acquiring reference material for quantification is costly and time-consuming.…
Authors not listed
Quantitative Structure Activity Relationship (QSAR) remains an effective tool for early-stage chemical modelling and virtual screening in drug design. The advancements in this field are led by two core paradigms, 1) descriptor engineering, where complex fixed-length vectors of compounds are generated and conventional…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Authors not listed
The integration of machine learning methods is transforming many areas of research by, for instance, accelerating molecular dynamics simulations and enabling improved prediction and optimization of chemical reactions. However, despite this progress, the adoption of data-driven approaches in atomic layer deposition…