21 papers · ranked by Valyu relevance
David Maxwell Chickering, David Heckerman
We describe two techniques that significantly improve the running time of several standard machine-learning algorithms when data is sparse. The first technique is an algorithm that efficiently extracts one-way and two-way counts—either real or expected from discrete data. Extracting such counts is a fundamental step in…
Niloufar Dousti Mousavi, Hani Aldirawi, Jie Yang, Jari Louhelainen
Categorical data analysis becomes challenging when high-dimensional sparse covariates are involved, which is often the case for omics data. We introduce a statistical procedure based on multinomial logistic regression analysis for such scenarios, including variable screening, model selection, order selection for…
Vartan Choulakian
Visualization and interpretation of contingency tables by correspondence analysis (CA), as developed by Benz´ecri, have a rich structure based on Euclidean geometry. However, it is a well established fact that, often CA is very sensitive to sparse contingency tables, where we characterize sparsity as the existence of…
Muhammad Taimoor Khan, Anila Usman
Sparse storage formats are techniques for storing and processing the sparse matrix data efficiently. The performance of these storage formats depend upon the distribution of non-zeros, within the matrix in different dimensions. In order to have better results we need a technique that suits best the organization of data…
Syed Muhammad Atif, Anees Ahmed, Sameer Qazi
Modern smart distribution system requires storage, transmission and processing of big data generated by sensors installed in electric meters. On one hand, this data is essentially required for intelligent decision making by smart grid but on the other hand storage, transmission and processing of that huge amount of…
S. Park, E. Ceulemans, K. Van Deun
Principal component analysis (PCA) is an important tool for analyzing large collections of variables. It functions both as a pre-processing tool to summarize many variables into components and as a method to reveal structure in data. Different coefficients play a central role in these two uses. One focuses on the…
Ravi Ganti, Rebecca Willett
This paper proposes a fast and accurate method for sparse regression in the presence of missing data. The underlying statistical model encapsulates the low-dimensional structure of the incomplete data matrix and the sparsity of the regression coefficients, and the proposed algorithm jointly learns the low-dimensional…
Ruilin Li, Christopher Chang, Yosuke Tanigawa, Balasubramanian Narasimhan + 3 more
We develop two efficient solvers for optimization problems arising from large-scale regularized regressions on millions of genetic variants sequenced from hundreds of thousands of individuals. These genetic variants are encoded by the values in the set {0, 1, 2, NA}. We take advantage of this fact and use two bits to…
Katrijn Van Deun, Tom F Wilderjans, Robert A van den Berg, Anestis Antoniadis + 1 more
'Anestis Antoniadis' 'Iven Van Mechelen'] 1 Background High throughput data are complex and methods that reveal structure underlying the data are most useful. Principal component analysis, frequently implemented as a singular value decomposition, is a popular technique in this respect. Nowadays often the challenge is…
Nan Miles Xi, Jingyi Jessica Li
Autoencoders are the backbones of many imputation methods that aim to relieve the sparsity issue in single-cell RNA sequencing (scRNA-seq) data. The imputation performance of an autoencoder relies on both the neural network architecture and the hyperparameter choice. So far, literature in the single-cell field lacks a…
Sanjar Adilov
Machine learning models for molecular-property prediction typically work with molecular representations in the form of fingerprints, descriptors, or graphs. In case of fingerprints and descriptors, molecular representations usually comprise thousands of features, which causes the curse of dimensionality for many…
Jean Daunizeau
So-called sparse estimators arise in the context of model fitting, when one a priori assumes that only a few (unknown) model parameters deviate from zero (Li, 2007). Typically, sparsity constraints can be useful when the estimation problem is under-determined, i.e. when number of parameters to estimate ( n ) is much…
Xinyu Zhou, Pengtao Dang, Haixu Tang, Laura Xianlu Peng + 6 more
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general framework that estimates…
Kelsey Hatzell, Yanjie Zheng
X-ray Computed Tomography (CT) is a non-invasive, non-destructive approach to imaging materials, material systems and engineered components in two- and three- dimensions. Acquisition of 3D images requires the collection of hundreds or thousands of through-thickness X-ray radiographic images from different angles. Such…
Sergio Hernández-Galaz, Ignacio Pezoa-Soto, Andrés Hernández-Oliveras, Sofía Rodriguez + 2 more
In Single-cell RNA-seq, observed zeroes are the mix between biological absence and technical limitations. However, current evaluation metrics fail to distinguish between these two states, focusing on reconstruction accuracy rather than the biological validity of edits. We introduce SPARE, a partition-aware framework…
David Buterez, Jon Paul Janet, Steven Kiddle, Pietro Liò
We investigate the potential of graph neural networks for transfer learning and improving molecular property prediction on sparse and expensive to acquire high-fidelity data by leveraging low-fidelity measurements as an inexpensive proxy for a targeted property ofinterest. This problem arises in discovery processes…
Matthias Scheffler, Stefan Bauer, Peter Benner, Tristan Bereau + 57 more
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Stamatia Zavitsanou, Zonghua Bo, Emanuele Casali, Matthew Langton + 1 more
Machine learning (ML) is currently transforming the field of chemistry by offering unparalleled efficiency in addressing complex challenges. Despite the progress made, a notable gap persists in the availability of user-friendly tools tailored to chemical problems involving small and sparse datasets. Here, we introduce…
Matthias Scheffler
Matthias Scheffler 1 , Stefan Bauer 2 , Peter Benner 3 , Tristan Bereau 4 , Volker Blum 5 , Mario Boley 6 , Christian Carbogno 7 , C. Richard A. Catlow 8 , Gerhard Dehm 9 , Sebastian Eibl 10 , Ralph Ernstorfer 11 , Ádám Fekete 12 , Lucas Foppa 1 , Peter Fratzl 13 , Christoph Freysoldt 9 , Baptiste Gault 9 , Luca M.…
Authors not listed
Identifying molecular structure based on spectroscopic readings is a key task in a va- riety of chemical and biological applications. Common spectroscopy techniques, such as Infrared (IR) Spectroscopy and Mass Spectrometry (MS), provide detailed information on the structure of molecular compounds but nonetheless…
Ryan Quey, Matthew A. Schiefer, Anmol Kiran, Bhavesh Patel
This manuscript provides the methods and outcomes of KnowMore, the Grand Prize winning automated knowledge discovery tool developed by our team during the 2021 NIH SPARC FAIR Data Codeathon. The National Institutes of Health Stimulating Peripheral Activity to Relieve Conditions (NIH SPARC) program generates rich…