26 papers · ranked by Valyu relevance
Yangruibo Ding, Jinjun Peng, Marcus J. Min, Gail E. Kaiser + 2 more
'Junfeng Yang' 'Baishakhi Ray'] Code Large Language Models (Code LLMs) have excelled at tasks like code completion but often miss deeper semantics such as execution effects and dynamic states. This paper aims to bridge the gap between Code LLMs' reliance on static text data and the need for thorough semantic…
Gangtao Xin, Pingyi Fan, Khaled B. Letaief, Mateu Sbert
In recent years, semantic communication has received significant attention from both academia and industry, driven by the growing demands for ultra-low latency and high-throughput capabilities in emerging intelligent services. Nonetheless, a comprehensive and effective theoretical framework for semantic communication…
Zijian Liang, Kai Niu, Jin Xu, Ping Zhang + 1 more
Recent semantic communication methods explore effective ways to expand the communication paradigm and improve the performance of communication systems. Nonetheless, a common problem with these methods is that the essence of semantics is not explicitly pointed out and directly utilized. A new epistemology suggests that…
Shi Yuxuan, Shuo Shao, Yongpeng Wu
A novel distributed source coding model which named semantic-aware multi-terminal source coding problem is proposed and studied in the paper. This is motivated by the new communication paradigm being aware of semantic information, in which invisible semantic features are observed by multiple agents, and both semantic…
Monoshi Kumar Roy, Simin Chen, Benjamin Steenhoek, Jinjun Peng + 3 more
Understanding and reasoning about code semantics is essential for enhancing code LLMs' abilities to solve real-world software engineering (SE) tasks. Although several code reasoning benchmarks exist, most rely on synthetic datasets or educational coding problems and focus on coarse-grained reasoning tasks such as…
Yue Cao, Youlong Wu, Lixiang Lian, Meixia Tao + 1 more
This study proposes a separate source-channel coding (SSCC) framework to address semantic communication challenges in MIMO systems, overcoming the limitations of joint source-channel coding (JSCC) in channel adaptation and model reusability. Traditional systems suffer from bit-level redundancy in 6G, while JSCC…
Melissa Franch, Elizabeth A. Mickiewicz, James L. Belanger, Assia Chericoni + 10 more
As we listen to speech, our brains actively compute the meaning of individual words. Inspired by the success of large language models (LLMs), we hypothesized that the brain employs vectorial coding principles, such that meaning is reflected in distributed activity of single neurons. We recorded responses of hundreds of…
Patrick Thomson, Rob Rix, Nicolas Wu, Tom Schrijvers
GitHub hosts hundreds of millions of code repositories written in hundreds of different programming languages. In addition to its hosting services, GitHub provides data and insights into code, such as vulnerability analysis and code navigation, with which users can improve and understand their software development…
Adam Štorek, Mukur Gupta, Samira Hajizadeh, Prashast Srivastava + 1 more
'Suman Jana'] Although modern Large Language Models (LLMs) support extremely large contexts, their effectiveness in utilizing long context for code reasoning remains unclear. This paper investigates LLM reasoning ability over code snippets within large repositories and how it relates to their recall ability.…
Roychoudhury, Abhik
At the code level, common software tasks include code generation, testing, and program repair. Design level software tasks may include architecture exploration, requirements understanding, and requirements enforcement at the code level. Each of these software tasks involves micro-decisions which can be taken…
Yun-Fei Liu, Marina Bedny
Programming is a cornerstone of modern society, yet its cognitive and neural basis remains poorly understood. In this study, we test the hypothesis that programming “recycles” pre-existing neural mechanisms and representations in fronto-parietal reasoning networks. Using fMRI, we scanned programming-naïve…
Ahmed Abdu, Zhengjun Zhai, Hakim A. Abdo, Redhwan Algabri + 3 more
'Mohammed A. Al-masni' 'Mannan Saeed Muhammad' 'Yeong Hyeon Gu'] Software defect prediction aims to find a reliable method for predicting defects in a particular software project and assisting software engineers in allocating limited resources to release high-quality software products. While most earlier research has…
Pankaj Thorat, Adnan Qidwai, Adrija Dhar, Aishwariya Chakraborty + 3 more
'Anand Eswaran' 'Hima Patel' 'Praveen Jayachandran'] Data profiling, in the context of machine learning, is the process of examining and analyzing data to create useful statistics. These statistics are used both as an aid for better comprehension of the properties of data as well as for a variety of downstream data…
Authors not listed
Bayesian optimization (BO) has become increasingly important for experimental optimization across scientific domains, yet implementing BO pipelines requires significant programming expertise and familiarity with specialized frameworks. This creates a barrier for domain experts who could benefit from BO but lack the…
Kun Sun
Expectation and memory have been found to play crucial roles in human language comprehension. Currently, the effects of both expectation and memory can be estimated using computational methods. Computational metrics of surprisal and semantic relevance, which represent expectation and memory respectively, have been…
Cristian Robledo, Francesca Sallicati, Gaël de Chalendar, Marcos Fernández + 4 more
'Marcos Fernández' 'Pablo de Castro' 'Eduardo Martín' 'Javier Gutiérrez' 'Yannis Bouachera'] This paper aims to introduce the innovative work carried out in the Horizon 2020 DECODER project - acronym for “DEveloper COmpanion for Documented and annotatEd code Reference” - (Grant Agreement no. 824231) by linking the…
Yuchen Wang, Shangxin Guo, Chee Wei Tan
Local-Cloud Copilot Framework Authors: ['Yuchen Wang' 'Shangxin Guo' 'Chee Wei Tan'] Abstract—The advancements in cloud-based Large Languages Models (LLMs) have revolutionized AI-assisted programming. However, their integration into certain local development environments like ones within the Apple software ecosystem…
Authors not listed
The nanosafety domain has seen significant advancements in data generation and sharing, yet challenges remain in ensuring data interoperability and reuse. This article focuses on developing a semantic interoperability framework for nanosafety data to maximize the FAIRness (Findability, Accessibility, Interoperability…
Authors not listed
While virtual libraries of synthetically accessible compounds have exploded in size to many billions, our capacity to extract valuable drug leads from these vast databases remains limited by computational resources. To overcome this, we developed SLICE SMARTS and Logic In ChEmistry), a powerful new tool designed for…
Pieter Floris Jacobs, Robert Pollice
Scientists across domains are often challenged to master domain-specific languages (DSLs) for their research, which are merely a means to an end but are pervasive in fields like computational chemistry. Automated code generation promises to overcome this barrier, allowing researchers to focus on their core expertise.…
Nazia Bibi, Tauseef Rana, Ayesha Maqbool, Farkhanda Afzal + 3 more
'Ali Akgül' 'Manuel De la Sen' 'Rebeca P. Díaz\xa0Redondo'] The development of robotic applications necessitates the availability of useful, adaptable, and accessible programming frameworks. Robotic, IoT, and sensor-based systems open up new possibilities for the development of innovative applications, taking advantage…
Meng Yang, Haiping Huang, Lichao Huang, Nan Zhang + 3 more
Interpretation of non-coding genome remains an unsolved challenge in human genetics due to impracticality of exhaustively annotate biochemically active elements in all conditions. Deep learning based computational approaches emerge recently to help interpretating non-coding regions. Here we present LOGO (Language of…
Rachana Niranjan Murthy, Sai Teja Potu, Akhil Thomas, Lokesh Mishra + 2 more
Retrieving structured materials information from unstructured textual data is essential for data mining and automatically developing comprehensive ontologies. Information extraction is a complex task composed of multiple subtasks and thus often relies on systems of task-specialized language models. A foundation…
Authors not listed
The scarcity and expense of fatigue data limits optimal design of components and constrains companies to a few well qualified materials when safety-critical applications are concerned. This research investigates different strategies to improve extraction of structured information from unstructured scientific…
Robert Haase, Christian Tischer, Jean-Karim Hériché, Nico Scherf
In the computational age, life-scientists often have to write Python code to solve bio-image analysis (BIA) problems. Many of them have not been formally trained in programming though. Code-generation, or coding assistance in general, with Large Language Models (LLMs) can have a clear impact on BIA. To the best of our…
Andrea Gurioli, Maurizio Gabbrielli, Stefano Zacchiroli, Stefan Wagner
'Stefan Wagner'] Code stylometry is the application of stylometry techniques to determine the authorship of software source code snippets. It is used in the industry to address use cases like plagiarism detection, code audits, and code review assignments. Most works in the code stylometry literature use machine…