24 papers · ranked by Valyu relevance
Daniel S. Roll, Zeyneb Kurt, Yulei Li, Wai Lok Woo + 1 more
This preliminary study covers the construction and application of a Graph-based Retrieval-Augmented Generation (GraphRAG) system integrating a multimodal LLM, Large Language and Vision Assistant (LLaVA) with graph database software (Neo4j) to enhance LLM output quality through structured knowledge retrieval. This is…
Edmund Evangelista, Fathima Ruba, Salman Bukhari, Amril Nazir + 2 more
Background Gestational diabetes mellitus (GDM) is a prevalent chronic condition that affects maternal and fetal health outcomes worldwide, increasingly in underserved populations. While generative artificial intelligence (AI) and large language models (LLMs) have shown promise in health care, their application in GDM…
Yukun Cao, Zhiqiang Gao, Zhiyang Li, Xike Xie + 1 more
for Design Space Exploration Authors: ['Yukun Cao' 'Zhiqiang Gao' 'Zhiyang Li' 'Xike Xie' 'S Kevin Zhou'] GraphRAG integrates (knowledge) graphs with large language models (LLMs) to improve reasoning accuracy and contextual relevance. Despite its promising applications and strong relevance to multiple research…
Renjie Liu, Haitian Jiang, Xiao Yan, Bo Tang + 1 more
GraphRAG enhances large language models (LLMs) to generate quality answers for user questions by retrieving related facts from external knowledge graphs. Existing GraphRAG methods adopt a fixed graph traversal strategy for fact retrieval but we observe that user questions come in different types and require different…
Zhiyi Xiang, Chuanjie Wu, Qinggang Zhang, Shengyuan Chen + 4 more
Graph retrieval-augmented generation (GraphRAG) has emerged as a powerful paradigm for enhancing large language models (LLMs) with external knowledge. It leverages graphs to model the hierarchical structure between specific concepts, enabling more coherent and effective knowledge retrieval for accurate reasoning.…
Jie Song, Jinhua Feng, Yuxin Zhang, Cheng Bi + 12 more
Personalized perioperative fluid therapy is important for reducing postoperative complications and adverse outcomes. Although large language models (LLMs) show promise in healthcare, their application in fluid therapy remains challenged by hallucinations, limited domain-specific knowledge, and insufficient…
Çerağ Oğuztüzün, Zhenxiang Gao, Rong Xu
Title: Summary: One of the primary challenges in biomedical research is the interpretation of complex genomic relationships and the prediction of functional interactions across the genome. Tokenvizz is a novel tool for genomic analysis that enhances data discovery and visualization by combining GraphRAG-inspired…
Kai Guo, Harry Shomer, Shenglai Zeng, Haoyu Han + 2 more
'Jiliang Tang'] In recent years, large language models (LLMs) have revolutionized the field of natural language processing. However, they often suffer from knowledge gaps and hallucinations. Graph retrieval-augmented generation (GraphRAG) enhances LLM reasoning by integrating structured knowledge from external graphs.…
Haoyu Han, Harry Shomer, Yu Wang, Lei Yuan + 5 more
'Bo Long' 'Hui Liu' 'Jiliang Tang'] Retrieval-Augmented Generation (RAG) enhances the performance of LLMs across various tasks by retrieving relevant information from external sources, particularly on textbased data. For structured data, such as knowledge graphs, GraphRAG has been widely used to retrieve relevant…
Çerağ Oğuztüzün, Zhenxiang Gao, Rong Xu
One of the primary challenges in biomedical research is the interpretation of complex genomic relationships and the prediction of functional interactions across the genome. Tokenvizz is a novel tool for genomic analysis that enhances data discovery and visualization by combining GraphRAG-inspired tokenization with…
Herbert George, Gowtham Baratam, DS Dhyaneesh
Drug repurposing has become a crucial strategy to accelerate drug discovery and reduce development costs. Conventional drug development is time consuming and expensive; it often takes more than a decade and billions of dollars to bring a new drug into the market. To address these challenges, this work puts forward a…
Nina Marthe, Matthias Zytnicki, Francois Sabot
The increasing availability of genome sequences has highlighted the limitations of using a single reference genome to represent the diversity within a species. Pangenomes, encompassing the genomic information from multiple genomes, offer thus a more comprehensive representation of intraspecific diversity. However…
Davide Torre, Davide Chicco, Jacqui Chetty
The challenge of analyzing high-dimensional data affects many scientific disciplines, from pharmacology to chemistry and biology. Traditional dimensionality reduction methods often oversimplify data, making it difficult to interpret individual points. This distortion can complicate the visualization of mutual distances…
Jon Mitchell Ambler, Shandukani Mulaudzi, Nicola Mulder
As sequencing technology improves, the concept of a single reference genome is becoming increasingly restricting. In the case of Mycobacterium tuberculosis, one must often choose between using a genome that is closely related to the isolate, or one that is annotated in detail. One promising solution to this problem is…
Heming Zhang, Shunning Liang, Tim Xu, Wenyu Li + 15 more
Artificial intelligence (AI) is revolutionizing scientific discovery because of its super capability, following the neural scaling laws, to integrate and analyze large-scale datasets to mine knowledge. Foundation models, large language models (LLMs) and large vision models (LVMs), are among the most important…
Mehmet Aziz Yirik, Maria Sorokina, Christoph Steinbeck
The generation of constitutional isomer chemical spaces has been a subject of cheminformatics since the early 1960s, with applications in structure elucidation and elsewhere. In order to perform such a generation efficiently, exhaustively and isomorphism-free, the structure generator needs to ensure the building of…
Caitlin C. Bannan, David Mobley
Force fields are used in a variety of research fields including computer-aided drug design, biomaterials, and polymer chemistry. However, force fields also continue to limit the accuracy of predictions of physical properties. Current parameterization of these force fields involves a huge amount of human effort -- often…
Jordan M. Eizenga, Adam M. Novak, Emily Kobayashi, Flavia Villani + 6 more
Pangenomics is a growing field within computational genomics. Many pangenomic analyses use bidirected sequence graphs as their core data model. However, implementing and correctly using this data model can be difficult, and the scale of pangenomic data sets can be challenging to work at. These challenges have impeded…
Authors not listed
Molecular property prediction has become essential in accelerating advancements in drug discovery and materials science. Graph Neural Networks have recently demonstrated remarkable success in molecular representation learning; however, their broader adoption is impeded by two significant challenges: (1) data scarcity…
Authors not listed
Curried functions provide a systematic way of transforming multi-argument functions into nested singleargument functions. This transformation allows partial application and supports many central principles of functional programming. Their extension, called curried 𝑘-ary functions, naturally generalizes the familiar…
Authors not listed
Liquid chromatography (LC) is a cornerstone of analytical separations, but comparing the retention times (RTs) for different LC methods is difficult because of variations in experimental parameters such as column type and solvent gradient. Nevertheless, RTs are powerful metrics in tandem mass spectrometry (MS2) that…
Anna Lisiecka, Agnieszka Kowalewska, Norbert Dojer
Pangenome graphs conveniently represent genetic variation within a population. Several types of such graphs have been proposed, with varying properties and potential applications. Among them, variation graphs (VGs) seem best suited to replace reference genomes in sequencing data processing, while whole genome…
Authors not listed
We present a unified, set–theoretic framework that extends molecular graphs to hypergraphs and superhypergraphs via iterated power sets. We define Molecular Graphs, Molecular HyperGraphs, and Molecular SuperHyperGraphs, and develop four complements over them: Weighted, Rough, Neural, and Multipolar frameworks. We prove…
David Buterez, Jon Paul Janet, Steven Kiddle, Pietro Liò
We investigate the potential of graph neural networks for transfer learning and improving molecular property prediction on sparse and expensive to acquire high-fidelity data by leveraging low-fidelity measurements as an inexpensive proxy for a targeted property ofinterest. This problem arises in discovery processes…