23 papers · ranked by Valyu relevance
Phil Ostheimer, Mayank Nagda, Marius Kloft, Sophie Fellenz
Sparse data is ubiquitous, appearing in numerous domains, from economics and recommender systems to astronomy and biomedical sciences. However, efficiently and realistically generating sparse data remains a significant challenge. We introduce Sparse Data Diffusion (SDD), a novel method for generating sparse data. SDD…
Skyler Ruiter, Seth Wolfgang, Marc Tunnell, Timothy J. Triche + 2 more
'Erin Carrier' 'Zachary J. DeBruine'] Compressed Sparse Column (CSC) and Coordinate (COO) are popular compression formats for sparse matrices. However, both CSC and COO are general purpose and cannot take advantage of any of the properties of the data other than sparsity, such as data redundancy. Highly redundant…
Julio Candanedo
Matrices and more generally multidimensional arrays, form the backbone of computational studies. In this paper we demonstrate increases in computational efficiency by performing partial-tracing/tensor-contractions on sparse-arrays. It was shown that sparse-arrays are really 3 dense-arrays (dense-shape, index-array, and…
Niloufar Dousti Mousavi, Hani Aldirawi, Jie Yang, Jari Louhelainen
Categorical data analysis becomes challenging when high-dimensional sparse covariates are involved, which is often the case for omics data. We introduce a statistical procedure based on multinomial logistic regression analysis for such scenarios, including variable screening, model selection, order selection for…
Maliheh Miri, Mohammad Taghi Sadeghi, Vahid Abootalebi
Sparse representation of signals has achieved satisfactory results in classification applications compared to the conventional methods. Microarray data, which are obtained from monitoring the expression levels of thousands of genes simultaneously, have very high dimensions in relation to the small number of samples.…
Edward Raff, Amol Khanna, Fred Lu
To the best of our knowledge, there are no methods today for training differentially private regression models on sparse input data. To remedy this, we adapt the Frank-Wolfe algorithm for L1 penalized linear regression to be aware of sparse inputs and to use them effectively. In doing so, we reduce the training time of…
S. Park, E. Ceulemans, K. Van Deun
Principal component analysis (PCA) is an important tool for analyzing large collections of variables. It functions both as a pre-processing tool to summarize many variables into components and as a method to reveal structure in data. Different coefficients play a central role in these two uses. One focuses on the…
Vishwas Choudhary, Binay Gupta, Anirban Chatterjee, Subhradip Paul + 2 more
'Kunal Banerjee' 'Vijay Srinivas Agneeswaran'] Missing values, widely called as sparsity in literature, is a common characteristic of many real-world datasets. Many imputation methods have been proposed to address this problem of data incompleteness or sparsity. However, the accuracy of a data imputation method for a…
Keisuke Teramoto, Kei Hirose
Summary. In the field of materials science and engineering, statistical analysis and machine learning techniques have recently been used to predict multiple material properties from an experimental design. These material properties correspond to response variables in the multivariate regression model. This study…
Sanjar Adilov
Machine learning models for molecular-property prediction typically work with molecular representations in the form of fingerprints, descriptors, or graphs. In case of fingerprints and descriptors, molecular representations usually comprise thousands of features, which causes the curse of dimensionality for many…
Kateřina Tschernosterová, Eva Trávníčková, Florencia Grattarola, Clara Rosse + 1 more
'Clara Rosse' 'Petr Keil'] Title: Abstract Here, we introduce SPARSE (acronym for "SPecies AcRoss ScalEs"), a simple and portable template for databases that can store data on species composition derived from ecological inventories, surveys and checklists, with emphasis on metadata describing sampling effort and…
Nan Miles Xi, Jingyi Jessica Li
Autoencoders are the backbones of many imputation methods that aim to relieve the sparsity issue in single-cell RNA sequencing (scRNA-seq) data. The imputation performance of an autoencoder relies on both the neural network architecture and the hyperparameter choice. So far, literature in the single-cell field lacks a…
Spencer L. Stahl, Stuart I. Benton
large unsteady datasets Authors: ['Spencer L. Stahl' 'Stuart I. Benton'] The cost of writing, transferring, and storing large amounts of data from unsteady simulations limits the accessibility of the entire solution, often leaving the majority of the flow under-sampled or not analyzed. For example, modeling the…
Rongbo Chen, Haojun Sun, Lifei Chen, Jianfei Zhang + 1 more
Markov models are extensively used for categorical sequence clustering and classification due to their inherent ability to capture complex chronological dependencies hidden in sequential data. Existing Markov models are based on an implicit assumption that the probability of the next state depends on the preceding…
Xinyu Zhou, Pengtao Dang, Haixu Tang, Laura Xianlu Peng + 6 more
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general framework that estimates…
Ju‐Chi Yu, Julie Le Borgne, Anjali Krishnan, Arnaud Gloaguen + 4 more
Generalized Singular Value Decomposition Authors: ['Ju‐Chi Yu' 'Julie Le Borgne' 'Anjali Krishnan' 'Arnaud Gloaguen' 'Cheng‐Ta Yang' 'Laura A. Rabin' 'Hervé Abdi' 'Vincent Guillemot'] Correspondence analysis, multiple correspondence analysis and their discriminant counterparts (i.e., discriminant simple correspondence…
Kelsey Hatzell, Yanjie Zheng
X-ray Computed Tomography (CT) is a non-invasive, non-destructive approach to imaging materials, material systems and engineered components in two- and three- dimensions. Acquisition of 3D images requires the collection of hundreds or thousands of through-thickness X-ray radiographic images from different angles. Such…
Sergio Hernández-Galaz, Ignacio Pezoa-Soto, Andrés Hernández-Oliveras, Sofía Rodriguez + 2 more
In Single-cell RNA-seq, observed zeroes are the mix between biological absence and technical limitations. However, current evaluation metrics fail to distinguish between these two states, focusing on reconstruction accuracy rather than the biological validity of edits. We introduce SPARE, a partition-aware framework…
James W. Webber, Kevin M. Elias
High dimensional transcriptome profiling, whether through next generation sequencing techniques or high-throughput arrays, may result in scattered variables with missing data. Data imputation is a common strategy to maximize the inclusion of samples by using statistical techniques to fill in missing values. However…
Sterling Baird, Jason R. Hall, Taylor D. Sparks
Would you rather search for a line inside a cube or a point inside a square? Physics-based simulations and wet-lab experiments often have symmetries (degeneracies) that allow reducing problem dimensionality or search space, but constraining these degeneracies is often unsupported or difficult to implement in many…
Matthias Scheffler, Stefan Bauer, Peter Benner, Tristan Bereau + 57 more
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Stamatia Zavitsanou, Zonghua Bo, Emanuele Casali, Matthew Langton + 1 more
Machine learning (ML) is currently transforming the field of chemistry by offering unparalleled efficiency in addressing complex challenges. Despite the progress made, a notable gap persists in the availability of user-friendly tools tailored to chemical problems involving small and sparse datasets. Here, we introduce…
Matthias Scheffler
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…