13 papers · ranked by Valyu relevance
Anastasia Drozdova, Ekaterina Trofimova, Polina Guseva, Anna Scherbakova + 2 more
The use of program code as a data source is increasingly expanding among data scientists. The purpose of the usage varies from the semantic classification of code to the automatic generation of programs. However, the machine learning model application is somewhat limited without annotating the code snippets. To address…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Ekaterina Trofimova, Emil Sataev, Andrey Ustyuzhanin, Xiangjie Kong
In the ever-evolving landscape of machine learning, seamless translation of natural language descriptions into executable code remains a formidable challenge. This article introduces Linguacodus, an innovative framework designed to tackle this challenge by deploying a dynamic pipeline that iteratively transforms…
Sarah E. Lindley, Yiyang Lu, Diwakar Shukla
Guide to Machine Learning for Small Molecule Design Authors: ['Sarah\nE. Lindley' 'Yiyang Lu' 'Diwakar Shukla'] Initially part of the field of artificial intelligence, machine learning (ML) has become a booming research area since branching out into its own field in the 1990s. After three decades of refinement, ML…
Nouf Alturayeif, Jameleddine Hassine, Bilal Alatas
With the increasing reliance on machine learning (ML) across diverse disciplines, ML code has been subject to a number of issues that impact its quality, such as lack of documentation, algorithmic biases, overfitting, lack of reproducibility, inadequate data preprocessing, and potential for data leakage, all of which…
Leo A. Celi, Luca Citi, Marzyeh Ghassemi, Tom J. Pollard + 1 more
'Leonie Anna Mueck'] Recent years have seen a surge of studies in machine learning in health and biomedicine, driven by digitalization of healthcare environments and increasingly accessible computer systems for conducting analyses. Many of us believe that these developments will lead to significant improvements in…
Rana Sandouka, Hamoud Aljamaan, Stephen Piccolo
Code smells are poor code design or implementation that affect the code maintenance process and reduce the software quality. Therefore, code smell detection is important in software building. Recent studies utilized machine learning algorithms for code smell detection. However, most of these studies focused on code…
Fabiano Pecorelli, Savanna Lujan, Valentina Lenarduzzi, Fabio Palomba + 1 more
Code smells are poor implementation choices that developers apply while evolving source code and that affect program maintainability. Multiple automated code smell detectors have been proposed: while most of them relied on heuristics applied over software metrics, a recent trend concerns the definition of machine…
Hamoud Aljamaan, Muhammad Aleem
Code smells refer to poor design and implementation choices by software engineers that might affect the overall software quality. Code smells detection using machine learning models has become a popular area to build effective models that are capable of detecting different code smells in multiple programming languages.…
Man-Fai Wong, Shangxin Guo, Ching-Nam Hang, Siu-Wai Ho + 2 more
'Chee-Wei Tan' 'Lei Wang'] This paper provides a comprehensive review of the literature concerning the utilization of Natural Language Processing (NLP) techniques, with a particular focus on transformer-based large language models (LLMs) trained using Big Code, within the domain of AI-assisted programming tasks. LLMs…
Yan Wang, Peng Jia, Luping Liu, Cheng Huang + 2 more
'Tao Song'] Security vulnerabilities play a vital role in network security system. Fuzzing technology is widely used as a vulnerability discovery technology to reduce damage in advance. However, traditional fuzz testing faces many challenges, such as how to mutate input seed files, how to increase code coverage, and…
Absalom E. Ezugwu, Olaide N. Oyelade, Abiodun M. Ikotun, Jeffery O. Agushaka + 1 more
The machine learning (ML) paradigm has gained much popularity today. Its algorithmic models are employed in every field, such as natural language processing, pattern recognition, object detection, image recognition, earth observation and many other research areas. In fact, machine learning technologies and their…
Sanaa Kaddoura, Daniela Elena Popescu, Jude D. Hemanth, Vimal Shanmuganathan
'Vimal Shanmuganathan'] Examinations or assessments play a vital role in every student’s life; they determine their future and career paths. The COVID pandemic has left adverse impacts in all areas, including the academic field. The regularized classroom learning and face-to-face real-time examinations were not…