21 papers · ranked by Valyu relevance
Tingting Mu
Matrix factorisation is a fundamental tool for exploiting low-dimensional structure in high-dimensional data, with applications such as data compression, denoising, structure discovery, interpretable representation learning, and dimensionality reduction. Compared to conventional two-factor models, matrix…
Mahbod Nouri, David Rotermund, Alberto Garcia-Ortiz, Klaus R. Pawelzik
Considering biological constraints in artificial neural networks has led to dramatic improvements in performance. Nevertheless, to date, the positivity of long-range signals in the cortex has not been shown to yield improvements. While Non-negative matrix factorization (NMF) captures biological constraints of positive…
Volkan Sevinç, Nikolas Kontemeniotis, Theodoros Perdikis, Michail Tsagris
Non--negative matrix factorization (NMF) has become an established dimensionality reduction technique for extracting latent structures from non--negative data and has found widespread applications in fields such as bioinformatics, text mining, image analysis, and recommender systems. As the popularity of NMF has…
Ko Abe, Shintaro Yuki, Teppei Shimamura
Combinatorial indexing-based single-cell RNA sequencing methods such as sci-RNA-seq and sci-RNA-seq3 now enable the profiling of millions of cells, producing expression matrices that are both extremely sparse and high-dimensional. Conventional nonnegative matrix factorization (NMF) provides an interpretable framework…
Cindy Fang, Kelsey D. Montgomery, Sarah E. Maguire, Anthony D. Ramnauth + 8 more
Recent advances in spatially-resolved transcriptomics have enabled profiling of gene expression in a spatial context, which has led to the generation of large-scale single-cell and spatial atlases with computationally-derived cell type or spatial domain labels. An increasingly important task with these data has become…
Ibrahim, Mubaraka Sani, Isah Charles Saidu, Lehel Csató
The growing popularity of group activities increased the need to develop methods for providing recommendations to a group of users based on the collective preferences of the group members. Several group recommender systems have been proposed, but these methods often struggle due to sparsity and high-dimensionality of…
Yu-Ting Liu, Timothy J Triche, Zachary J DeBruine
Large single-cell atlases now span tens of millions of cells, yet few provide reusable and interpretable reference representations that support direct biological reasoning at atlas-scale. Here, we present an interpretable Non-negative Matrix Factorization reference embedding of 28.5 million healthy cells and…
Lara Kassab, Erin George, Deanna Needell, Haowen Geng + 2 more
There has been a recent critical need to study fairness and bias in machine learning (ML) algorithms. Since there is clearly no one-size-fits-all solution to fairness, ML methods should be developed alongside bias mitigation strategies that are practical and approachable to the practitioner. Motivated by recent work on…
Yudong Wei, Liang Zhang, Bingcong Li, Niao He
Low-rank matrix optimization is often carried out via the Burer-Monteiro (BM) formulation, but choosing the factorization rank $r$ is delicate and can substantially slow optimization. We propose a unified framework, termed direction-magnitude decomposition (DMD), that decomposes the optimization variable to improve…
Jin Deng, Junjie Lan, Ruolan Du, Tao Xu + 4 more
The high recurrence rate of tumor limits the growth of precision medicine, whereas the exploration of correlations in multimodal data enables mining of features linked to tumor recurrence, ultimately identifying prospective biomarkers. Nevertheless, existing multimodal approaches centered on genetic molecular data…
Geert Roelof van der Ploeg, Fred T. G. White, Rasmus Riemer Jakobsen, Johan A. Westerhuis + 3 more
The rapid growth of high-dimensional biological data has necessitated advanced data fusion techniques to integrate and interpret complex multi-omics and longitudinal datasets. Shared and unshared structure across such datasets can be identified in an unsupervised manner with Advanced Coupled Matrix and Tensor…
Cristian Castiglione, Alexandre Segers, Lieven Clement, Davide Risso
Title: Summary Single-cell RNA sequencing allows the quantification of gene expression at the individual cell level, enabling the study of cellular heterogeneity and gene expression dynamics. Dimensionality reduction is a common preprocessing step critical for the visualization, clustering, and phenotypic…
Eric Weine, Peter Carbonetto, Rafael A. Irizarry, Matthew Stephens
Poisson non-negative matrix factorization (NMF) is a widely used method to find interpretable "parts-based" decompositions of count data. While many variants of Poisson NMF exist, existing methods assume that the "parts" in the decomposition combine additively. This assumption may be natural in some settings, but not…
Tim Faverjon, Jean-Philippe Cointet, Pedro Ramaciotti, Shuai Liu
Recommendations play a crucial role in shaping informational diets on social media, raising concerns regarding potential consequences such as political segregation. We take an algorithm explanability approach, as opposed to a description of recommendations, to show how recommenders inadvertently create geometrical…
SeungJoo Lee, Yong-Chan Park, U. Kang, George Vousden
How can we accurately decompose a temporal irregular tensor along while incorporating a related knowledge graph tensor in both offline and online streaming settings? PARAFAC2 decomposition is widely applied to the analysis of irregular tensors consisting of matrices with varying row sizes. In both offline and online…
Rui Zhang, Jinhang Liu, Wenbo Zhang
High-dimensional and incomplete (HDI) data are prevalent in many real-world big data scenarios. Latent factor models serve as a common representation learning approach, capable of uncovering informative latent factors from such data. Nevertheless, most existing latent factor models rely solely on gradient descent for…
Xiaoge Zhang, Zhengyu Fang, Kaiyu Tang, Huiyuan Chen + 1 more
Targeted drug therapies offer a promising approach for treating complex diseases, with combinational drug therapies often employed to enhance therapeutic efficacy. However, unintended drug-drug interactions may undermine treatment outcomes or cause adverse side effects. In this work, we propose a novel joint learning…
Authors not listed
Developing a transferable classical force field (FF) has historically been a lengthy, expert-informed process. In this work, we integrate optimization, machine learning, and data science techniques to accelerate the systematic design and parameterization of transferable FF models. As a demonstration, we create…
Authors not listed
Data-driven approaches offer great potential for accelerating ab initio electronic structure calculations of molecules and materials but their transferability is often limited due to the vast amount of data needed for training, including when addressing the need to fine-tune universal models for each specific system to…
Authors not listed
The discovery of chemically novel or structurally anomalous metal-organic frameworks (MOFs) is essential for expanding reticular design space and enhancing dataset reliability. We present CHEM-AD (Chemically Unusual Metal–organic Frameworks via Autoencoder-based Detection), a label-free, CPU-efficient pipeline that…
Karim G. Habashy, Benjamin D. Evans, Dan F. M. Goodman, Jeffrey S. Bowers
The genomic mechanisms that efficiently encode the initial architecture and synaptic connectivity of neural circuits remain poorly understood. We hypothesise that two primary mechanisms — spatial encoding and factorisation — enable a limited genome to initialise networks of billions of neurons. Spatial encoding, a form…