24 papers · ranked by Valyu relevance
Samir Char, Nathaniel Corley, Sarah Alamdari, Kevin K. Yang + 1 more
Understanding the protein sequence-function relationship is essential for advancing protein biology and engineering. However, fewer than 1% of known protein sequences have human-verified functions. While deep learning methods have demonstrated promise for protein function prediction, current models are limited to…
Zhuoming Liu, Xuefeng Hu, Ram Nevatia
We propose a new setting for detecting unseen objects called Zero-shot Annotation object Detection (ZAD). It expands the zero-shot object detection setting by allowing the novel objects to exist in the training images and restricts the additional information the detector uses to novel category names. Recently, to…
He Zhang, Xiaolong Fu
Models: A Multi-Class and Multi-Frame Approach in DailyLife Authors: ['He Zhang' 'Xiaolong Fu'] This study investigates the feasibility and performance of using large language models (LLMs) to automatically annotate human emotions in everyday scenarios. We conducted experiments on the DailyLife subset of the publicly…
Samir Char, Nathaniel Corley, Sarah Alamdari, Kevin K Yang + 2 more
Based on these observations, we hypothesized that similar protein sequences would tend to have similar function labels and that ProtNote protein embeddings from the sequence projection head would reflect this relationship. To investigate this, we first define two ways to capture the similarity between a pair of…
Meysam Alizadeh, Maël Kubli, Zeynab Samei, Shirin Dehghani + 4 more
This paper studies the performance of open-source Large Language Models (LLMs) in text classification tasks typical for political science research. By examining tasks like stance, topic, and relevance classification, we aim to guide scholars in making informed decisions about their use of LLMs for text analysis and to…
Oscar Sainz, Iker García-Ferrero, Rodrigo Agerri, Oier López de Lacalle + 2 more
'Oier López de Lacalle' 'Germán Rigau' 'Eneko Agirre'] Large Language Models (LLMs) combined with instruction tuning have made significant progress when generalizing to unseen tasks. However, they have been less successful in Information Extraction (IE), lagging behind task-specific models. Typically, IE tasks are…
Yanwei Fu, Timothy M. Hospedales, Tao Xiang, Shaogang Gong
—Most existing zero-shot learning approaches exploit transfer learning via an intermediate semantic representation shared between an annotated auxiliary dataset and a target dataset with different classes and no annotation. A projection from a low-level feature space to the semantic representation space is learned from…
Jaromir Savelka, Kevin D. Ashley
The emergence of ChatGPT has sensitized the general public, including the legal profession, to large language models' (LLMs) potential uses (e.g., document drafting, question answering, and summarization). Although recent studies have shown how well the technology performs in diverse semantic annotation tasks focused…
Angelo Basile, Marc Franco-Salvador, Paolo Rosso
Zero-shot text classifiers based on label descriptions embed an input text and a set of labels into the same space: measures such as cosine similarity can then be used to select the most similar label description to the input text as the predicted label. In a true zero-shot setup, designing good label descriptions is…
Aitor González-Marfil, Estibaliz Gómez-de-Mariscal, Ignacio Arganda-Carreras
We present DINOSim, a novel approach leveraging the DINOv2 pretrained encoder for zero-shot object detection and segmentation in electron microscopy datasets. By exploiting semantic embeddings, DINOSim generates pseudo-labels from patch distances to a user-selected reference, which are subsequently employed in a…
Zhangxuan Gu, Siyuan Zhou, Li Niu, Zihan Zhao + 1 more
Semantic segmentation, aiming at classifying each pixel in one image, heavily relies on the dense pixel-wise annotations [5, 25, 26, 38, 50, 53]. To reduce the annotation pressure, leveraging weak annotations like image-level [34, 35, 47], box-level [18, 36], or scribblelevel [24] annotations for semantic segmentation…
Xueyi Zhong, Liye Zhao, Licheng Peng, Guodong Yang + 2 more
Relation extraction serves as an essential task for knowledge acquisition and management, defined as determining the relation between two annotated entities from a piece of text. Over recent years, zero-shot learning has been introduced to train relation extraction models due to the expensive cost of incessantly…
Zihan Ye, Shreyank N Gowda, Shiming Chen, Yaochu Jin + 2 more
'Xiaobo Jin'] Zero-shot learning (ZSL) aims to recognize unseen classes by aligning images with intermediate class semantics, like human-annotated concepts or class definitions. An emerging alternative leverages Large-scale Language Models (LLMs) to automatically generate class documents. However, these methods often…
Matthew D. Turner, Abhishek Appaji, Nibras Ar Rakib, Pedram Golnari + 6 more
We show that recent (mid-to-late 2024) commercial large language models (LLMs) are capable of good quality metadata extraction and annotation with very little work on the part of investigators for several exemplar real-world annotation tasks in the neuroimaging literature. We investigated the GPT-4o LLM from OpenAI…
Beibei Yu, Cheng Xie, Peng Tang, Bin Li + 1 more
Almost all existing zero-shot learning methods work only on benchmark datasets (e.g., CUB, SUN, AwA, FLO and aPY) which have already provided pre-defined attributes for all the classes. These methods thus are hard to apply on real-world datasets (like ImageNet) since there are no such pre-defined attributes in the data…
Authors not listed
Molecular property prediction is a fundamental task in computational chemistry with critical applications in drug discovery and materials science. While recent works have explored Large Language Models (LLMs) for this task, they primarily rely on textual molecular representations such as SMILES/SELFIES, which can be…
Jiongxin Liu, Jiameng Le, Chuanru Wei, Mingming Liu + 1 more
Predicting drug–target interactions (DTI) for entirely unseen drugs or proteins—the cold-start problem—remains a critical challenge in computational drug discovery. While sequence-based methods naturally support zero-shot generalization, they often ignore relational topology, and existing graph-based approaches either…
Emily Allaway, Kathleen McKeown
A major challenge in stance detection is the large (potentially infinite) and diverse set of stance topics. Collecting data for such a set is unrealistic due to both the expense of annotation and the continuous creation of new real-world topics (e.g., a new politician runs for office). Furthermore, stancetaking occurs…
David Kainer
Ontologies are highly prevalent in biology and medicine and are always evolving. Annotating biological text, such as observed phenotype descriptions, with ontology terms is a challenging and tedious task. The process of annotation requires a contextual understanding of the input text and of the ontological terms…
Hao Yuan, Parker Hicks, Mansooreh Ahmadian, Kayla Johnson + 2 more
Reusing massive collections of publicly available biomedical data can significantly impact knowledge discovery. However, these public samples and studies are typically described using unstructured plain text, hindering the findability and further reuse of the data. To combat this problem, we propose txt2onto 2.0, a…
Rachana Niranjan Murthy, Sai Teja Potu, Akhil Thomas, Lokesh Mishra + 2 more
Retrieving structured materials information from unstructured textual data is essential for data mining and automatically developing comprehensive ontologies. Information extraction is a complex task composed of multiple subtasks and thus often relies on systems of task-specialized language models. A foundation…
Melanie Vollmar, Santosh Tirunagari, Deborah Harrus, David Armstrong + 5 more
We present a novel system that leverages curators in the loop to develop a dataset and model for detecting residue-level functional annotations and other protein structure features from standard publication text. Our approach involves the integration of data from multiple resources, including PDBe, EuropePMC…
Authors not listed
The scarcity and expense of fatigue data limits optimal design of components and constrains companies to a few well qualified materials when safety-critical applications are concerned. This research investigates different strategies to improve extraction of structured information from unstructured scientific…
Authors not listed
Computational models predicting the sites of metabolism (SOM) of small or- ganic molecules have become invaluable tools for studying and optimizing the metabolic properties of xenobiotics. However, the performance of SOM predic- tors has shown signs of plateauing in recent years, primarily due to the limited…