14 papers · ranked by Valyu relevance
Guangshuo Cao, Yi Shen, Jianghong Wu, Haoyu Chao + 2 more
We present CellReasoner, a lightweight, open-source large language model (LLM) tailored for single-cell type annotation. We introduced a compact training strategy that activates the reasoning capabilities of 7B-parameter LLMs using only 380 high-quality chain-of-thought exemplars. CellReasoner directly maps cell-level…
Yiran Song, Muyao Tang, Qi Liu, Haofei Wang + 3 more
Cell type annotation is critical for interpreting single-cell transcriptomic data but remains challenging due to uncertain cellular clustering granularity and inconsistent labeling across studies. Here we present GPTAnno, an automated, ontology-tree-guided, uncertainty-aware, hierarchical cell type annotation method…
Felix Fischer, David S. Fischer, Evan Biederstedt, Alexandra-Chloé Villani + 1 more
Identifying cellular identities (both novel and well-studied) is one of the key use cases in single-cell transcriptomics. While supervised machine learning has been leveraged to automate cell annotation predictions for some time, there has been relatively little progress both in scaling neural networks to large data…
Sebastiano Cultrera di Montesano, Davide D’Ascenzo, Srivatsan Raghavan, Ava P. Amini + 2 more
Accurately annotating cell types is essential for extracting biological insight from single-cell RNA-seq data. Although cell types are naturally organized into hierarchical ontologies, most computational models do not explicitly incorporate this structure into their training objectives. We introduce a hierarchical…
Luni Hu, Qianqian Chen, Ping Qiu, Hua Qin + 7 more
Standardizing cell type annotations across single-cell RNA-seq datasets remains a major challenge due to inconsistencies in nomenclature, variation in annotation granularity, and the presence of rare or previously unseen populations. We present UniCell, a hierarchical annotation framework that combines Cell Ontology…
Arman Kazmi, Deepshikha Singh, Shashank Jatav, Soumya Luthra
Recent research has shown the impressive capability of large language models like GPT-4 in various downstream tasks in single-cell data analysis. Among these tasks, cell type annotation remains particularly challenging, with researchers exploring various methods to improve accuracy and efficiency. While recent studies…
Dimitrios Kleftogiannnis, Sonia Gavasso, Benedicte Sjo Tislevoll, Nisha van der Meer + 9 more
Mass cytometry by time-of-flight (CyTOF) is an emerging technology allowing for in-depth characterisation of cellular heterogeneity in cancer and other diseases. However, computational identification of cell populations from CyTOF, and utilisation of single cell data for biomarker discoveries faces several technical…
Stephen R. Williams, Fedor Grab, Govinda M. Kamath, Yerdos Ordabayev + 13 more
Cell type annotation in single-cell RNA sequencing (scRNA-seq) experiments is the fundamental step of assigning cell types to individual cells or clusters of cells based on their gene expression profiles. This process is crucial for developing biological insights from scRNA-seq experiments. We present a service that…
Rufus H. Daw, Harry R. Deijnen, Magnus Rattray, John R. Grainger
Single-cell RNA sequencing (scRNA-seq) cell annotation techniques rely on the matching of known defining marker genes to a given cell population. However, these methods may lack robustness to dynamic fluctuations in cell marker expression between patients, samples and pathologies. The advent of easy-to-implement…
George Crowley, Stephen R. Quake
We developed an open-source package called AnnDictionary (https://github.com/ggit12/anndictionary/) to facilitate the parallel, independent analysis of multiple anndata. AnnDictionary is built on top of LangChain and AnnData and supports all common large language model (LLM) providers. AnnDictionary only requires 1…
Aanchal Mongia, Diane C. Saunders, Yue J. Wang, Marcela Brissova + 6 more
Cellular composition and anatomical organization influence normal and aberrant organ functions. Emerging spatial single-cell proteomic assays such as Image Mass Cytometry (IMC) and Co-Detection by Indexing (CODEX) have facilitated the study of cellular composition and organization by enabling high-throughput…
Dezheng Han, Yibin Jia, Ruxiao Chen, Wenjie Han + 2 more
To enable precise and fully automated cell type annotation with large language models (LLMs), we developed a graph-structured feature–marker database to retrieve entities linked to differential genes for cell reconstruction. We further designed a multi-task workflow to optimize the annotation process. Compared to…
Haohuai He, Zhenchao Tang, Guanxing Chen, Fan Xu + 6 more
Single-cell analysis has revolutionized our understanding of cellular heterogeneity, yet current approaches face challenges in efficiency and interpretability. In this study, we present scKAN, a framework that leverages Kolmogorov-Arnold Networks for interpretable single-cell analysis through three key innovations…
Jonathan Karin, Reshef Mintz, Barak Raveh, Mor Nitzan
Single-cell and spatial genomics datasets can be organized and interpreted by annotating single cells to distinct types, states, locations, or phenotypes. However, cell annotations are inherently ambiguous, as discrete labels with subjective interpretations are assigned to heterogeneous cell populations based on noisy…