23 papers · ranked by Valyu relevance
Arpit Merchant, Ananth Mahadevan, Michael Mathioudakis, Adam Lipowski
'Adam Lipowski'] The task of node classification concerns a network where nodes are associated with labels, but labels are known only for some of the nodes. The task consists of inferring the unknown labels given the known node labels, the structure of the network, and other known node attributes. Common node…
Wei Xiao-wen, Weiwei Liu, Yibing Zhan, Bo Du + 1 more
Node classification is a fundamental graph-based task that aims to predict the classes of unlabeled nodes, for which Graph Neural Networks (GNNs) are the state-of-the-art methods. Current GNNs assume that nodes in the training set contribute equally during training. However, the quality of training nodes varies…
Aleksandar Tomčić, Miloš Savić, Miloš Radovanović
In the last two decades we are witnessing a huge increase of valuable big data structured in the form of graphs or networks. To apply traditional machine learning and data analytic techniques to such data it is necessary to transform graphs into vector-based representations that preserve the most essential structural…
Xuan Wu, Yifei Shen, Fangzhou Ge, Caihua Shan + 3 more
'Xiangguo Sun' 'Hong Cheng'] Node classification is a fundamental task in graph analysis, with broad applications across various fields. Recent breakthroughs in Large Language Models (LLMs) have enabled LLM-based approaches for this task. Although many studies demonstrate the impressive performance of LLM-based…
Zhenpeng Liu, Shengcong Zhang, Jialiang Zhang, Mingxiao Jiang + 2 more
'Yi Liu' 'Alessandro Pluchino'] Most Heterogeneous Information Network (HIN) embedding methods use meta-paths to guide random walks to sample from HIN and perform representation learning in order to overcome the bias of traditional random walks that are more biased towards high-order nodes. Their performance depends on…
F. Zeng, Wensheng Gan, Philip S. Yu
—The class imbalance problem refers to the disproportionate distribution of samples across different classes within a dataset, where the minority classes are significantly underrepresented. This issue is also prevalent in graph-structured data. Most graph neural networks (GNNs) implicitly assume a balanced class…
Zeng, Fanlong, Gan, Wensheng + 4 more
The problem of class imbalance refers to an uneven distribution of quantity among classes in a dataset, where some classes are significantly underrepresented compared to others. Class imbalance is also prevalent in graph-structured data. Graph neural networks (GNNs) are typically based on the assumption of class…
Masoud Kargar, Nasim Jelodari, Alireza Assadzadeh
graphs and graph convolutional networks for high-level feature extraction Authors: ['Masoud Kargar' 'Nasim Jelodari' 'Alireza Assadzadeh'] Graphs, comprising nodes and edges, visually depict relationships and structures, posing challenges in extracting high-level features due to their intricate connections Multiple…
Nan Chen, Zemin Liu, Bryan Hooi, Bingsheng He + 2 more
In real-world applications, node classification on graphs often faces the challenge of class imbalance, where majority classes dominate training, resulting in biased model performance. Traditional GNNs often struggle in such scenarios, as they tend to overfit to majority classes while underrepresenting minority…
Khaoula Ait Rai, Mustapha Machkour, Jilali Antari
Researchers have paid a lot of attention to complex networks in recent decades. Due to their rapid evolution, they turn into a major scientific and innovative field. Several studies on complex networks are carried out, and other subjects are evolving every day such as the challenge of detecting influential nodes. In…
Siyu Tao, Yang Yang, Xin Liu, Yimiao Feng + 1 more
Predicting biomedical interactions is crucial for understanding various biological processes and drug discovery. Graph neural networks (GNNs) are promising in identifying novel interactions when extensive labeled data are available. However, labeling biomedical interactions is often time-consuming and labor-intensive…
Lida Kanari, Stanislav Schmidt, Francesco Casalegno, Emilie Delattre + 7 more
The shape of neuronal morphologies plays a critical role in determining their dynamical properties and the functionality of the brain. With an abundance of neuronal morphology reconstructions, a robust definition of cell types is important to understand their role in brain functionality. However, an objective…
Liuhai Wang, Xin Du, Bo Jiang, Weifeng Pan + 3 more
'Dongsheng Liu' 'Amelia Carolina Sparavigna'] Software maintenance is indispensable in the software development process. Developers need to spend a lot of time and energy to understand the software when maintaining the software, which increases the difficulty of software maintenance. It is a feasible method to…
Paul Chon, William B. Andreopoulos
Protein language models (PLMs) are shown to be powerful predictors of protein structure and function but their internal mechanisms remain poorly understood. Recent mechanistic interpretability methods have decomposed PLM representations into interpretable features, but they have not combined methods on a single…
Yihan Deng, Kerstin Denecke
The Swiss classification of surgical interventions (CHOP) has to be used in daily practice by physicians to classify clinical procedures. Its purpose is to encode the delivered healthcare services for the sake of quality assurance and billing. For encoding a procedure, a code of a maximal of 6-digits has to be selected…
Authors not listed
The identification of kinetically feasible reaction pathways that connect a reactant to its product, including numerous intermediates and transition states, is crucial for predicting chemical reactions and elucidating reaction mechanisms. However, as molecular systems become increasingly complex or larger, the number…
Lidong Fu, Xin Ma, Zengfa Dou, Yun Bai + 2 more
In the field of complex network analysis, accurately identifying key nodes is crucial for understanding and controlling information propagation. Although several local centrality methods have been proposed, their accuracy may be compromised if interactions between nodes and their neighbors are not fully considered. To…
Renming Liu, Matthew Hirn, Arjun Krishnan
Accurately representing biological networks in a low-dimensional space, also known as network embedding, is a critical step in network-based machine learning and is carried out widely using node2vec, an unsupervised method based on biased random walks. However, while many networks, including functional gene interaction…
Authors not listed
Chemical reactions typically follow mechanistic templates and hence fall into a manageable number of clearly distinguishable classes that usually labeled by names of chemists who discovered or explored them. These ``named reactions'' form the core of reaction ontologies and are associated with specific synthetic…
Paul Kruse, Caroline Ring
This paper presents the treecompareR package for R, which provides tools for reproducible visualizations of data through the use of taxonomies. The package builds on developments from ggplot2 and ggtree to provide visualizations tailored for use with taxonomic classification data. Additionally, it provides tools that…
Authors not listed
Natural language processing with the help of large language models such as ChatGPT has become ubiquitous in many software applications and allows users to interact even with complex hardware or software in an intuitive way. The recent concepts of Self-Driving Labs and Material Acceleration Platforms stand to benefit…
Authors not listed
Diversity and properties of ring systems contained in small molecules are of high interest for applications such as drug discovery or material sciences. In the present work we extract, analyse and classify ring systems found in small molecule compounds of open access databases such as PubChem, ChEMBL, DrugCentral…
Authors not listed
Today, machine learning models are employed extensively to predict the physicochemical and biological properties of molecules. Their performance is typically evaluated on in-distribution (ID) data, i.e., data originating from the same distribution as the training data. However, the real-world applications of such…