13 papers · ranked by Valyu relevance
Anastasia Drozdova, Ekaterina Trofimova, Polina Guseva, Anna Scherbakova + 2 more
The use of program code as a data source is increasingly expanding among data scientists. The purpose of the usage varies from the semantic classification of code to the automatic generation of programs. However, the machine learning model application is somewhat limited without annotating the code snippets. To address…
Valeriy Berezovskiy, Anastasia Gorodilova, Ekaterina Trofimova, Andrey Ustyuzhanin + 1 more
'Andrey Ustyuzhanin' 'Syed Hassan Shah'] Program code has recently become a valuable active data source for training various data science models, from code classification to controlled code synthesis. Annotating code snippets play an essential role in such tasks. This article presents a novel approach that leverages…
Nawaf Alomari, Amal Alazba, Hamoud Aljamaan, Mohammad Alshayeb
Context: Code smells indicate poor software design, affecting maintainability. Accurate detection is vital for refactoring and quality improvement. However, existing datasets often frame detection as single-label classification, limiting realism. Objective: This paper develops a multi-label dataset for code smell…
Ekaterina Trofimova, Emil Sataev, Andrey Ustyuzhanin, Xiangjie Kong
In the ever-evolving landscape of machine learning, seamless translation of natural language descriptions into executable code remains a formidable challenge. This article introduces Linguacodus, an innovative framework designed to tackle this challenge by deploying a dynamic pipeline that iteratively transforms…
Rana Sandouka, Hamoud Aljamaan, Stephen Piccolo
Code smells are poor code design or implementation that affect the code maintenance process and reduce the software quality. Therefore, code smell detection is important in software building. Recent studies utilized machine learning algorithms for code smell detection. However, most of these studies focused on code…
Dhivyabharathi Ramasamy, Cristina Sarasua, Alberto Bacchelli, Abraham Bernstein
'Abraham Bernstein'] Despite the ubiquity of data science, we are far from rigorously understanding how coding in data science is performed. Even though the scientific literature has hinted at the iterative and explorative nature of data science coding, we need further empirical evidence to understand this practice and…
Abdullah Al-Boghdady, Mohammad El-Ramly, Khaled Wassif
Internet of Things (IoT) 's devices are ubiquitous and operate in a heterogonous environment with potential security breaches. IoT Operating Systems (IoT OSs) are the backbone software for running such devices. If IoT OSs are vulnerable to security breaches, higher-level security measures may not help. This paper aims…
Hamoud Aljamaan, Muhammad Aleem
Code smells refer to poor design and implementation choices by software engineers that might affect the overall software quality. Code smells detection using machine learning models has become a popular area to build effective models that are capable of detecting different code smells in multiple programming languages.…
Nouf Alturayeif, Jameleddine Hassine, Bilal Alatas
With the increasing reliance on machine learning (ML) across diverse disciplines, ML code has been subject to a number of issues that impact its quality, such as lack of documentation, algorithmic biases, overfitting, lack of reproducibility, inadequate data preprocessing, and potential for data leakage, all of which…
Rajwant Singh Rao, Seema Dewangan, Alok Mishra, Manjari Gupta
Detecting code smells may be highly helpful for reducing maintenance costs and raising source code quality. Code smells facilitate developers or researchers to understand several types of design flaws. Code smells with high severity can cause significant problems for the software and may cause challenges for the…
Sabine Eichhorn, Franz Niklas Mitze, Fritz Wagner, Inga Marte Charlott Seuthe + 6 more
To verify medical representativeness, the most common ICD-10 codes and OPS codes of each database were selected. The results for each database were first ranked separately and then compared. The primary aim was to compare how many of the codes found in DESTATIS were also found in the ML-Network among the most frequent…
Elisabetta Manduchi, Joseph D. Romano, Jason H. Moore
The genetic analysis of complex traits has been dominated by parametric statistical methods due to their theoretical properties, ease of use, computational efficiency, and intuitive interpretation. However, there are likely to be patterns arising from complex genetic architectures which are more easily detected and…
Jacqueline A. Jansen, Artür Manukyan, Nour Al Khoury, Altuna Akalin + 1 more
'Diego A. Forero'] Data analysis is constrained by a shortage of skilled experts, particularly in biology, where detailed data analysis and subsequent interpretation is vital for understanding complex biological processes and developing new treatments and diagnostics. One possible solution to this shortage in experts…