28 papers · ranked by Valyu relevance
Alex Bronstein, Pablo Sprechmann, Guillermo Sapiro
We present a comprehensive framework for structured sparse coding and modeling extending the recent ideas of using learnable fast regressors to approximate exact sparse codes. For this purpose, we propose an efficient feed forward architecture derived from the iteration of the block-coordinate algorithm. This…
Benjamin Cox, Vı́ctor Elvira
—State-space models (SSMs) are a powerful statistical tool for modelling time-varying systems via a latent state. In these models, the latent state is never directly observed. Instead, a sequence of data points related to the state are obtained. The linear-Gaussian state-space model is widely used, since it allows for…
S. Park, E. Ceulemans, K. Van Deun
Principal component analysis (PCA) is an important tool for analyzing large collections of variables. It functions both as a pre-processing tool to summarize many variables into components and as a method to reveal structure in data. Different coefficients play a central role in these two uses. One focuses on the…
Andreas Christ Sølvsten Jørgensen, Marc Sturrock, Atiyo Ghosh, Vahid Shahrezaei
The reverse engineering of gene regulatory networks based on gene expression data is a challenging inference task. A related problem in computational systems biology lies in identifying signalling networks that perform particular functions, such as adaptation. Indeed, for many research questions, there is an ongoing…
Basile Jumentier, Kevin Caye, Barbara Heude, Johanna Lepeule + 1 more
Association of phenotypes or exposures with genomic and epigenomic data faces important statistical challenges. One of these challenges is to remove variation due to unobserved confounding factors, such as individual ancestry or cell-type composition in tissues. This issue can be addressed with penalized latent factor…
Gabriel Barello, Adam S. Charles, Jonathan W. Pillow
The sparse coding model posits that the visual system has evolved to efficiently code natural stimuli using a sparse set of features from an overcomplete dictionary. The classic sparse coding model suffers from two key limitations, however: (1) computing the neural response to an image patch requires minimizing a…
Anthony Christidis, Stefan Van Aelst, Ruben H. Zamar
The two primary approaches for high-dimensional regression problems are sparse methods (e.g. best subset selection which uses the ℓ0-norm in the penalty) and ensemble methods (e.g. random forests). Although sparse methods typically yield interpretable models, they are often outperformed in terms of prediction accuracy…
Armin Askari, Alexandre d’Aspremont, Laurent El Ghaoui
Due to its linear complexity, naive Bayes classification remains an attractive supervised learning method, especially in very large-scale settings. We propose a sparse version of naive Bayes, which can be used for feature selection. This leads to a combinatorial maximum-likelihood problem, for which we provide an exact…
Yanbo Lian, Anthony N. Burkitt
Sparse coding, predictive coding and divisive normalization have each been found to be principles that underlie the function of neural circuits in many parts of the brain, supported by substantial experimental evidence. However, the connections between these related principles are still poorly understood. In this…
Sanjar Adilov
Machine learning models for molecular-property prediction typically work with molecular representations in the form of fingerprints, descriptors, or graphs. In case of fingerprints and descriptors, molecular representations usually comprise thousands of features, which causes the curse of dimensionality for many…
José Camacho, Age K. Smilde, Edoardo Saccenti, Johan A. Westerhuis
Sparse Principal Component Analysis (sPCA) is a popular matrix factorization approach based on Principal Component Analysis (PCA) that combines variance maximization and sparsity with the ultimate goal of improving data interpretation. When moving from PCA to sPCA, there are a number of implications that the…
Barbara E. Engelhardt, Ryan P. Adams
Substantial research on structured sparsity has contributed to analysis of many different applications. However, there have been few Bayesian procedures among this work. Here, we develop a Bayesian model for structured sparsity that uses a Gaussian process (GP) to share parameters of the sparsity-inducing prior in…
Yanbo Lian, Anthony N. Burkitt, Boris S. Gutkin
Sparse coding, predictive coding and divisive normalization have each been found to be principles that underlie the function of neural circuits in many parts of the brain, supported by substantial experimental evidence. However, the connections between these related principles are still poorly understood. Sparse coding…
Katrijn Van Deun, Tom F Wilderjans, Robert A van den Berg, Anestis Antoniadis + 1 more
'Anestis Antoniadis' 'Iven Van Mechelen'] 1 Background High throughput data are complex and methods that reveal structure underlying the data are most useful. Principal component analysis, frequently implemented as a singular value decomposition, is a popular technique in this respect. Nowadays often the challenge is…
Donghwan Lee, Woojoo Lee, Youngjo Lee, Yudi Pawitan
Background Principal component analysis (PCA) has gained popularity as a method for the analysis of high-dimensional genomic data. However, it is often difficult to interpret the results because the principal components are linear combinations of all variables, and the coefficients (loadings) are typically nonzero.…
Zhengguo Gu, Katrijn Van Deun
This article introduces a package developed for R (R Core Team, [25]) for performing an integrated analysis of multiple data blocks (i.e., linked data) coming from different sources. The methods in this package combine simultaneous component analysis (SCA) with structured selection of variables. The key feature of this…
Xinyu Zhou, Pengtao Dang, Haixu Tang, Laura Xianlu Peng + 6 more
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general framework that estimates…
Katrijn Van Deun, Elise A. V. Crompvoets, Eva Ceulemans
Background Data analysis methods are usually subdivided in two distinct classes: There are methods for prediction and there are methods for exploration. In practice, however, there often is a need to learn from the data in both ways. For example, when predicting the antibody titers a few weeks after vaccination on the…
Xinyu Zhou, Pengtao Dang, Xiao Wang, Laura Xianlu Peng + 7 more
Spatial transcriptomics (ST) data demands models that recover how associations among molecular and cellular features change across tissue while contending with noise, collinearity, cell mixing, and thousands of predictors. We present Spatially Smooth Sparse Regression (S3R), a general statistical framework that…
Kun Du
Likelihood Authors: ['Kun Du'] This paper compares convex and non-convex penalized likelihood methods in high-dimensional statistical modeling, focusing on their strengths and limitations. Convex penalties, such as LASSO, offer computational efficiency and strong theoretical guarantees, but often introduce bias in…
Authors not listed
High-throughput experimentation (HTE) in materials science generates vast, high-dimensional datasets relating synthesis parameters to material properties. While machine learning (ML) models excel at predicting properties from these parameters, they often fail to distinguish causal drivers from merely correlated…
Chen Qu, Paul Houston, Qi Yu, Riccardo Conte + 3 more
Hamiltonian matrices in electronic and nuclear contexts are highly compute-intensive to calculate, mainly due to the cost for the potential matrix. Typically these matrices contain many off-diagonal elements that are orders of magnitude smaller than diagonal elements. We illustrate that here for vibrational H-matrices…
Authors not listed
Acoustic measurements of batteries are known to be correlated to their state-of-charge, creating opportunities for state estimation that do not rely on electrical signals. State estimators are typically parametric models fitted from data, often from the broad toolbox of machine learning. Such models can be easily…
Sterling Baird, Jason R. Hall, Taylor D. Sparks
Would you rather search for a line inside a cube or a point inside a square? Physics-based simulations and wet-lab experiments often have symmetries (degeneracies) that allow reducing problem dimensionality or search space, but constraining these degeneracies is often unsupported or difficult to implement in many…
Sterling Baird, Jason R. Hall, Taylor D. Sparks
Would you rather search for a line inside a cube or a point inside a square? This type of solution degeneracy often exists in physics-based simulations and wet-lab experiments, but constraining these degeneracies is often unsupported or difficult to implement in many optimization packages, requiring additional time and…
Authors not listed
Quantitative Structure-Activity Relationship (QSAR) modeling is a pillar of computational drug discovery. However, standard machine learning (ML) models are often confounded by the high-dimensional and intensely correlated nature of molecular descriptors. A model may identify a "bulk" property (e.g., molecular weight)…
David Buterez, Jon Paul Janet, Steven Kiddle, Pietro Liò
We investigate the potential of graph neural networks for transfer learning and improving molecular property prediction on sparse and expensive to acquire high-fidelity data by leveraging low-fidelity measurements as an inexpensive proxy for a targeted property ofinterest. This problem arises in discovery processes…
Michael Dumelle, Matt Higham, Jay M. Ver Hoef, A. K. M. Anisur Rahman
'A. K. M. Anisur Rahman'] spmodel is an R package used to fit, summarize, and predict for a variety spatial statistical models applied to point-referenced or areal (lattice) data. Parameters are estimated using various methods, including likelihood-based optimization and weighted least squares based on variograms.…