24 papers · ranked by Valyu relevance
Harvey, Ethan, Loevlie, Dennis Johan + 2 more
Multiple instance learning (MIL) is often used in medical imaging to classify high-resolution 2D images by processing patches or classify 3D volumes by processing slices. However, conventional MIL approaches treat instances separately, ignoring contextual relationships such as the appearance of nearby patches or slices…
Ehsan Ahmed Dhrubo, Mohammad Mahmudul Alam, Edward Raff, Tim Oates + 1 more
Multiple Instance Learning (MIL) tasks impose a strict logical constraint: a bag is labeled positive if and only if at least one instance within it is positive. While this iff constraint aligns with many real-world applications, recent work has shown that most deep learning-based MIL approaches violate it, leading to…
Christian Hallgrimson, Y. Lydia Li, Claire A. Shou, Ben Cardoen + 5 more
Single-molecule localization microscopy (SMLM) achieves nanoscale imaging of complex protein structures in the cell. However, the ability to capture structural variability across cell conditions (cell lines, gene expression, treatment) from 3D point cloud SMLM data remains limited. We present siMILe, a…
Alexander Möllers, Marvin Sextro, Julius Hense, Gabriel Dernbach + 1 more
Multiple Instance Learning (MIL) addresses problems where supervision is available at the level of bags of instances and has been successfully applied in fields ranging from computational pathology to satellite imagery. Nevertheless, existing algorithms struggle in the low-label regime that characterizes many…
Puneet Sharma, Kristian Hindberg, Eibe Frank, Benedicte Schelde‐Olesen + 1 more
Identifying unique polyps in colon capsule endoscopy (CCE) images is a critical yet challenging task for medical personnel due to the large volume of images, the cognitive load it creates for clinicians, and the ambiguity in labeling specific frames. This paper formulates this problem as a multi-instance learning (MIL)…
Marina D’Amato, Jeroen van der Laak, Francesco Ciompi
Accurate tumor detection in digital pathology whole-slide images (WSIs) is crucial for cancer diagnosis and treatment planning. Multiple Instance Learning (MIL) has emerged as a widely used approach for weakly-supervised tumor detection with large-scale data without the need for manual annotations. However, traditional…
Ethan Levien
In multiple instance regression (MIR) data are organized into bags (collections of instances in feature space) and the goal is to learn a mapping that assigns labels to bags. A typical assumption is that there is a so-called concept point in feature space, the proximity to which dictates the bag label. Motivated by…
Ziying Yang, Michael Baudis
Somatic copy number aberrations (CNAs) represent a distinct class of genomic mutations associated with oncogenetic effects. Over the past three decades, significant volumes of CNA data have been generated through molecular-cytogenetic and genome sequencing-based techniques. These data have been pivotal in identifying…
Zeyu Gao, Anyu Mao, Yuxing Dong, Hannah Clayton + 7 more
Spatial quantification is a critical step in most computational pathology tasks, from guiding pathologists to areas of clinical interest to discovering tissue phenotypes behind novel biomarkers. To circumvent the need for manual annotations, modern computational pathology methods have favored multiple-instance learning…
Simon Grouard, Christian Esposito, Jean El Khoury, Valérie Ducret + 11 more
Accurate prediction of patient outcomes remains a major challenge in oncology. While recent machine learning (ML) approaches often rely on bulk omics lacking spatial resolution or histology-based multiple instance learning (MIL), spatial transcriptomics (SpT) provides a unique opportunity to capture both molecular…
Christian Hallgrimson, Y. Lydia Li, Claire A. Shou, Ben Cardoen + 5 more
Single-molecule localization microscopy (SMLM) achieves nanoscale imaging of complex protein structures in the cell. However, the ability to capture structural variability across cell conditions (cell lines, gene expression, treatment) from 3D point cloud SMLM data remains limited. We present siMILe, a weakly…
Kyeonghun Jeong, Jinwook Choi, Kwangsoo Kim
Title: Summary Linking cellular states to clinical phenotypes is a major challenge in single-cell analysis. Here, we present single-cell multiple instance learning for sample classification and associated subpopulation discovery (scMILD), a weakly supervised multiple instance learning framework that robustly identifies…
Alexander Rakowski, Christoph Lippert, Heather J. Cordell
Identifying causal genetic variants in a computational manner remains an open problem. Training end-to-end prediction models is not possible without large ground-truth datasets, while results of genome-wide association studies (GWAS) are entangled by linkage disequilibrium (LD), and gene expression datasets do not…
Salome Kazeminia, Muhammed Furkan Dasdelen, Bastian Rieck, Carsten Marr
Microscopic images of cells and tissues are central to disease diagnosis. In computational pathology, multiple instance learning (MIL) has emerged as a key paradigm for analyzing numerous images within a single patient sample. While the representative distribution of cells in a sample is important for diagnosis…
Qinqin Xie, Yongheng Sun, Yuxia Liang, Yu Shang + 6 more
Background According to the 2021 WHO classification of tumors of the central nervous system, isocitrate dehydrogenase (IDH) status serve an independent prognostic biomarker and is closely associated with tumor diagnosis and treatment response. At present, the determination of IDH status still relies on invasive…
Authors not listed
Machine olfaction—the artificial replication of the sense of smell—faces significant challenges due to the absence of large, standardized training datasets. Unlike vision, language, and audio models, which benefit from extensive corpora such as ImageNet, GLUE, and AudioSet, olfaction lacks scaled equivalents and…
Bin Liu, Haoyu Peng, Zhijia Wei, Jiajing Zhang + 1 more
Batch selection is crucial for improving both training efficiency and predictive performance in deep multi-label classification (MLC). Existing batch selection methods typically rely on a single metric to assess instance importance and use static label weights to distinguish label significance, neglecting the dynamic…
Ibrahim Alsaggaf, Daniel Buchan, Cen Wan
Cell-type identification plays a fundamental role in single-cell RNA-Seq analytics. Thanks to the recent success of the contrastive learning paradigm, the accuracy of automatic cell-type identification has also been improved. In this work, we propose a novel contrastive learning-based cell-type identification method…
Authors not listed
Recent advances in machine learning force fields (MLFF) have significantly extended the reach of atomistic simulations. Continuous progress in this field requires reliable reference datasets, accurate MLFF architectures, and efficient active learning strategies to enable robust modeling of complex molecular and…
Eiram Mahera Sheikh, Alaa Tharwat, Constanze Schwan, Wolfram Schenck
Pretrained cell segmentation models have simplified and accelerated microscopy image analysis, but they often perform poorly on challenging datasets. Although these models can be adapted to new datasets with only a few annotated images, the effectiveness of fine-tuning depends critically on which images are selected…
Authors not listed
We present a new method for fingerprint- ing atomic configurations relevant to ML-IAM training and application, utilizing the ChIMES descriptor. These fingerprints enable rigor- ous analysis of statistical distinguishability be- tween configurations. Sample applications in- clude assessing diversity within ML-IAP…
Authors not listed
Quantitative Structure Activity Relationship (QSAR) remains an effective tool for early-stage chemical modelling and virtual screening in drug design. The advancements in this field are led by two core paradigms, 1) descriptor engineering, where complex fixed-length vectors of compounds are generated and conventional…
Authors not listed
Early prediction of drug-induced organ toxicity remains a major bottleneck in drug discovery and clinical pharmacotherapy. Most data-driven toxicity models behave as endpoint predictors: they output a label but provide limited transparency about why a compound is risky or which evidence channel dominated the decision.…
Authors not listed
Machine learning (ML) models are increasingly used in quantum chemistry, but their reliability hinges on uncertainty quantification (UQ). In this study, we compare two prominent UQ paradigms—Deep Evidential Regression (DER) and Deep Ensembles—on the QM9 and WS22 datasets, with a specific emphasis on the role of post…