15 papers · ranked by Valyu relevance
Anastasia Drozdova, Polina Guseva, Е. В. Трофимова, Anna Scherbakova + 1 more
'A. Ustyuzhanin'] Program code as a data source is gaining popularity in the data science community. Possible applications for models trained on such assets range from classification for data dimensionality reduction to automatic code generation. However, without annotation number of methods that could be applied is…
Ruibo Shi, Lili Tao, Rohan Saphal, Fran Silavong + 1 more
We present CV4Code, a compact and effective computer vision method for sourcecode understanding. Our method leverages the contextual and the structural information available from the code snippet by treating each snippet as a two-dimensional image, which naturally encodes the context and retains the underlying…
Syed Mehedi Hasan Nirob, Shamim Ehsan, Moqsadur Rahman, Summit Haque
—Large language models (LLMs) have made it remarkably easy to synthesize plausible source code from natural language prompts. While this accelerates software development and supports learning, it also raises new risks for academic integrity, authorship attribution, and responsible AI use. This paper investigates the…
Bart van Oort, Luís Cruz, Maurício Aniche, Arie van Deursen
—Artificial Intelligence (AI) and Machine Learning (ML) are pervasive in the current computer science landscape. Yet, there still exists a lack of software engineering experience and best practices in this field. One such best practice, static code analysis, can be used to find code smells, i.e., (potential) defects in…
Marc Oedingen, Raphael C. Engelhardt, Robin Denz, Maximilian Hammer + 1 more
In recent times, large language models (LLMs) have made significant strides in generating computer code, blurring the lines between code created by humans and code produced by artificial intelligence (AI). As these technologies evolve rapidly, it is crucial to explore how they influence code generation, especially…
Peter Hamfelt, Ricardo Britto, Lincoln S. Rocha, Camilo Almendra
Machine learning (ML) has rapidly grown in popularity, becoming vital to many industries. Currently, the research on code smells in ML applications lacks tools and studies that address the identi! cation and validity of ML-speci!c code smells. This work investigates suitable methods and tools to design and develop a…
Authors not listed
Accurate prediction of molecular properties is essential for computational design in many areas of chemistry. Deep learning has been used in these prediction tasks for a wide variety of molecular properties, and the availability of user-friendly, open-source software implementing such architectures has democratized…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Esben Bjerrum, Tobias Rastemo, Ross Irwin, Christos Kannas + 1 more
Recent years have seen a large interest in using the Simplified Molecular Input Line Entry System (SMILES) chemical language as input for deep learning architectures solving chemical tasks. Many successful applications have been demonstrated within de novo molecular design, quantitative structure-activity relationship…
Rıza Özçelik, Laura van Weesep, Sarah de Ruiter, Francesca Grisoni
In this work, we introduce peptidy -- a lightweight Python library that facilitates converting peptides (expressed as aminoacid sequences) to numerical representations suited to machine learning. peptidy is free from external dependencies, integrates seamlessly into modern Python environments, and supports a range of…
Niklas Kühl, Marc Goutier, Robin Hirt, Gerhard Satzger
The application of "machine learning" and "artificial intelligence" has become popular within the last decade. Both terms are frequently used in science and media, sometimes interchangeably, sometimes with different meanings. In this work, we aim to clarify the relationship between these terms and, in particular, to…
Muhammad Hanzla, Abdul Rehman Shinwari
Machine Learning (ML) can be defined as a class of Artificial Intelligence for automated data analysis, which is capable of detecting patterns in data. The extracted patterns can be used to predict un-known data or to assist in decision-making processes under uncertainty. Recent advances in experimental and…
Authors not listed
Recent advances in machine learning force fields (MLFF) have significantly extended the reach of atomistic simulations. Continuous progress in this field requires reliable reference datasets, accurate MLFF architectures, and efficient active learning strategies to enable robust modeling of complex molecular and…
Anubhav Jain
The number of studies that apply machine learning (ML) to materials science has been growing at a rate of approximately 1.67 times per year over the past decade. In this review, I examine this growth in various contexts. First, I present an analysis of the most commonly used tools (software, databases, materials…
Yannick Ureel, Maarten R. Dobbelaere, Yi Ouyang, Kevin De Ras + 3 more
By combining machine learning with design of experiments, so-called active machine learning, more efficient and cheaper research can be conducted. Machine learning algorithms are more flexible, and are better at investigating the processes spanning all length scales of chemical engineering. While the active machine…