21 papers · ranked by Valyu relevance
Luca Di Grazia, Michael Pradel
The immense amounts of source code provide ample challenges and opportunities during software development. To handle the size of code bases, developers commonly search for code, e.g., when trying to find where a particular feature is implemented or when looking for code examples to reuse. To support developers in…
Weisong Sun, Chunrong Fang, Yifei Ge, Yuling Hu + 5 more
'Quanjun Zhang' 'Xiuting Ge' 'Yang Liu' 'Zhenyu Chen'] WEISONG SUN, State Key Laboratory for Novel Software Technology, Nanjing University, China and School of Computer Science and Engineering, Nanyang Technological University, Singapore CHUNRONG FANG∗ , State Key Laboratory for Novel Software Technology, Nanjing…
Chao Liu, Xin Xia, David Lo, Cuiyun Gao + 2 more
Code search is a core software engineering task. Effective code search tools can help developers substantially improve their software development efficiency and effectiveness. In recent years, many code search studies have leveraged different techniques, such as deep learning and information retrieval approaches, to…
Chao Liu, Xin Xia, David Lo, Z. Liu + 2 more
To accelerate software development, developers frequently search and reuse existing code snippets from a large-scale codebase, e.g., GitHub. Over the years, researchers proposed many information retrieval (IR) based models for code search, but they fail to connect the semantic gap between query and code. An early…
Qihong Song, Haize Hu, Tebo Dai
Code search aims to search for code snippets from large codebase that are semantically related to natural query statements. Deep learning is a valuable method for solving code search tasks in which the quality of training data directly impacts the performance of deep-learning models. However, most existing…
Hao Wang, Jia Zhang, Yingce Xia, Jiang Bian + 2 more
'Tie‐Yan Liu'] Abstract—Semantic code search, which aims to retrieve code snippets relevant to a given natural language query, has attracted many research efforts with the purpose of accelerating software development. The huge amount of online publicly available code repositories has prompted the employment of deep…
Masudur Rahman, Jed Barson, Sydney Paul, Joshua Kayan + 5 more
'Federico Andrés Lois' 'Sebastián Fernandez Quezada' 'Christopher Parnin' 'Kathryn T. Stolee' 'Baishakhi Ray'] Search is an integral part of a software development process. Developers often use search engines to look for information during development, including reusable code snippets, API understanding, and reference…
Muhammad Hammad, Önder Babur, Hamid Abdul Basit, Mark van den Brand + 1 more
'Yilun Shang'] Software developers frequently reuse source code from repositories as it saves development time and effort. Code clones (similar code fragments) accumulated in these repositories represent often repeated functionalities and are candidates for reuse in an exploratory or rapid development. To facilitate…
Brooke Ballantyne Scott, Susan Baer, Ashley Farrell, Pat Lee + 3 more
Although libraries have provided online-mediated search services for more than forty years , there is not yet a practice standard to guide execution of searches, train searchers, or evaluate search performance. “Mediated search services” describe the search services offered by libraries and professional librarians…
Nazia Bibi, Tauseef Rana, Ayesha Maqbool, Farkhanda Afzal + 3 more
'Ali Akgül' 'Manuel De la Sen' 'Rebeca P. Díaz\xa0Redondo'] The development of robotic applications necessitates the availability of useful, adaptable, and accessible programming frameworks. Robotic, IoT, and sensor-based systems open up new possibilities for the development of innovative applications, taking advantage…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Yekaterina Shulgina, Sean R. Eddy
The genetic code has been proposed to be a “frozen accident”, but the discovery of alternative genetic codes over the past four decades has shown that it can evolve to some degree. Since most examples were found anecdotally, it is difficult to draw general conclusions about the evolutionary trajectories of codon…
Christian Lovis, Melissa Friesen, Carlos Lara, Hongchang Bao + 2 more
'Christopher J O Baker' 'Anil Adisesh'] Background In many research studies, the identification of social determinants is an important activity, in particular, information about occupations is frequently added to existing patient data. Such information is usually solicited during interviews with open-ended questions…
Genet Abay Shiferaw, Elien Vandermarliere, Niels Hulstaert, Ralf Gabriels + 2 more
Spectral similarity searching to identify peptide-derived MS/MS spectra is a promising technique, and different spectrum similarity search tools have therefore been developed. Each of these tools, however, comes with some limitations, mainly due to low processing speed and issues with handling large databases.…
Hamoud Aljamaan, Muhammad Aleem
Code smells refer to poor design and implementation choices by software engineers that might affect the overall software quality. Code smells detection using machine learning models has become a popular area to build effective models that are capable of detecting different code smells in multiple programming languages.…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Emilie S. Henault, Maria Harris Rasmussen, Jan H. Jensen
We attempt to explain why search algorithms can find molecules with particular properties in an enormous chemical space (ca 10 60 molecules) by considering only a tiny subset (typically 10 3−6 molecules). Using a very simple example, we show that the number of potential paths that the search algorithms can follow to…
Ash Sze, Soha Hassoun
Databases are indispensable in biological and biomedical research, hosting vast amounts of structured and unstructured data, facilitating the organization, retrieval, and analysis of complex data. Database access, however, remains a manual, tedious, and sometimes overwhelming, task. We investigate in this study the…
Hung Q. Vo, Huy Q. Vo, Son T. Ly, Zhihao Wan + 5 more
Conventional tissue image analysis software provides foundational capabilities for cellular analysis, including segmentation, basic morphological feature extraction, and spatial organization analysis. However, these tools often require manual intervention and are not well integrated with code-driven automation…
Zhi Ping, Haoling Zhang, Shihong Chen, Qianlong Zhuang + 2 more
Chamaeleo is currently the only collection library that focuses on adapting multiple well-established coding schemes for DNA storage. It provides a tool for researchers to study various coding schemes and apply them in practice. Chamaeleo adheres to the concept of high aggregation and low coupling for software design…
Huifang Ma, Zhicheng Ji
Large language models have shown remarkable capabilities in algorithm design, but their effectiveness in solving data science challenges remains poorly understood. We conducted a classroom experiment in which graduate students used large language models (LLMs) to solve biomedical data science challenges on Kaggle.…