Search · four archives
Search · four archives
23 papers · ranked by Valyu relevance
Steffen Grünewälder, Guy Lever, Luca Baldassarre, Sam Patterson + 2 more
'Arthur Gretton' 'Massimiliano Pontil'] We demonstrate an equivalence between reproducing kernel Hilbert space (RKHS) embeddings of conditional distributions and vector-valued regressors. This connection introduces a natural regularized loss function which the RKHS embeddings minimise, providing an intuitive…
Alicia Zeng, Jack Gallant
Encoding models based on word embeddings or artificial neural network (ANN) features reliably predict brain responses to naturalistic stimuli but remain difficult to interpret. A central limitation is superposition—the entanglement of distinct semantic features along correlated directions in dense embeddings, which…
Sha-Sha Wu, Mi-Xiao Hou, Chun-Mei Feng, Jin-Xing Liu
Feature selection and sample clustering play an important role in bioinformatics. Traditional feature selection methods separate sparse regression and embedding learning. Later, to effectively identify the significant features of the genomic data, Joint Embedding Learning and Sparse Regression (JELSR) is proposed.…
Kenneth L. Clarkson, David P. Woodruff
We design a new distribution over poly(rε−1 ) × n matrices S so that for any fixed n × d matrix A of rank r, with probability at least 9/10, kSAxk2 = (1 ± ε)kAxk2 simultaneously for all x ∈ R d . Such a matrix S is called a subspace embedding. Furthermore, SA can be computed in O(nnz(A))time, where nnz(A) is the number…
Krishnakumar Balasubramanian, Kai Yu, Guy Lebanon
We propose and analyze a novel framework for learning sparse representations, based on two statistical techniques: kernel smoothing and marginal regression. The proposed approach provides a flexible framework for incorporating feature similarity or temporal information present in data sets, via non-parametric kernel…
Shabarish Chenakkod, Michał Dereziński, Xiaoyu Dong, Mark Rudelson
It is known that the embedding dimension of an OSE must satisfy m ≥ d, and for any θ > 0, a Gaussian embedding matrix with m ≥ (1 + θ)d is an OSE with ϵ = Oθ(1). However, such optimal embedding dimension is not known for other embeddings. Of particular interest are sparse OSEs, having s ≪ m non-zeros per column…
Xue Yu, Yifan Sun, Hai-Jun Zhou
High-dimensional linear regression model is the most popular statistical model for high-dimensional data, but it is quite a challenging task to achieve a sparse set of regression coefficients. In this paper, we propose a simple heuristic algorithm to construct sparse high-dimensional linear regression models, which is…
Basile Jumentier, Kevin Caye, Barbara Heude, Johanna Lepeule + 1 more
Association of phenotypes or exposures with genomic and epigenomic data faces important statistical challenges. One of these challenges is to remove variation due to unobserved confounding factors, such as individual ancestry or cell-type composition in tissues. This issue can be addressed with penalized latent factor…
Yanjun Li, Bihan Wen, Hao Cheng, Yoram Bresler
Low-dimensional embeddings for data from disparate sources play critical roles in multi-modal machine learning, multimedia information retrieval, and bioinformatics. In this paper, we propose a supervised dimensionality reduction method that learns linear embeddings jointly for two feature vectors representing data of…
Junyang Qian, Yosuke Tanigawa, Ruilin Li, Robert Tibshirani + 2 more
In high-dimensional regression problems, often a relatively small subset of the features are relevant for predicting the outcome, and methods that impose sparsity on the solution are popular. When multiple correlated outcomes are available (multitask), reduced rank regression is an effective way to borrow strength and…
Rosember Guerra-Urzola, Katrijn Van Deun, Juan C. Vera, Klaas Sijtsma
'Klaas Sijtsma'] PCA is a popular tool for exploring and summarizing multivariate data, especially those consisting of many variables. PCA, however, is often not simple to interpret, as the components are a linear combination of the variables. To address this issue, numerous methods have been proposed to sparsify the…
Sanjar Adilov
Machine learning models for molecular-property prediction typically work with molecular representations in the form of fingerprints, descriptors, or graphs. In case of fingerprints and descriptors, molecular representations usually comprise thousands of features, which causes the curse of dimensionality for many…
Anwar O. Nunez-Elizalde, Alexander G. Huth, Jack L. Gallant
Predictive models for neural or fMRI data are often fit using regression methods that employ priors on the model parameters. One widely used method is ridge regression, which employs a spherical Gaussian prior that assumes equal and independent variance for all parameters. However, a spherical prior is not always…
Tan Guo, Xiaoheng Tan, Lei Zhang, Chaochen Xie + 1 more
Recently, low-rank and sparse model-based dimensionality reduction (DR) methods have aroused lots of interest. In this paper, we propose an effective supervised DR technique named block-diagonal constrained low-rank and sparse-based embedding (BLSE). BLSE has two steps, i.e., block-diagonal constrained low-rank and…
Bo Liu, Sanfeng Chen, Shuai Li, Yongsheng Liang
In this paper a new framework, called Compressive Kernelized Reinforcement Learning (CKRL), for computing near-optimal policies in sequential decision making with uncertainty is proposed via incorporating the non-adaptive data-independent Random Projections and nonparametric Kernelized Least-squares Policy Iteration…
Rui Meng, Herbert K. H. Lee, Braden Soper, Priyadip Ray
Gaussian processes are a flexible Bayesian nonparametric modelling approach that has been widely applied, but poses computational challenges. To address the poor scaling of exact inference methods, approximation methods based on sparse Gaussian processes (SGP) are attractive. An issue faced by SGP, especially in latent…
Shuang Li, Bing Liu, Chen Zhang
Traditional multiple kernel dimensionality reduction models are generally based on graph embedding and manifold assumption. But such assumption might be invalid for some high-dimensional or sparse data due to the curse of dimensionality, which has a negative influence on the performance of multiple kernel learning. In…
S. Park, E. Ceulemans, K. Van Deun
Principal component analysis (PCA) is an important tool for analyzing large collections of variables. It functions both as a pre-processing tool to summarize many variables into components and as a method to reveal structure in data. Different coefficients play a central role in these two uses. One focuses on the…
Authors not listed
Early-stage drug discovery often suffers from data scarcity and out-of-distribution (OOD) shifts, which constrain the reliability of predictive models. While deep learning has advanced representation learning from molecular and biological data, tabular modeling remains indispensable, particularly in small-sample and…
Poorya Parvizi, Francisco Azuaje, Evropi Theodoratou, Saturnino Luz
A network embedding approach reduces the analysis complexity of large biological networks by converting them to lowdimensional vector representations (features/embeddings). These lower-dimensional vectors can then be used in machine learning prediction tasks with a wide range of applications in computational biology…
Hasan M. Sayeed, Sterling G. Baird, Taylor D. Sparks
Capturing structure-property relationships of materials for property prediction using machine learning requires the representation or featurization of the structural aspects of materials at different levels, including atomic, crystal, and microscales. While crystal structure-based modeling techniques are effective for…
Xiaoyu Li, Fangfang Zhu, Wenwen Min
The rapid development of spatially resolved transcriptomics (SRT) technologies has provided unprecedented opportunities for exploring the structure of specific organs or tissues. However, these techniques (such as image-based SRT) can achieve single-cell resolution, but can only capture the expression levels of tens to…
Authors not listed
Real-world datasets in chemical engineering and bioengineering processes--such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials--can often be unlabelled or disorganized, rendering the training of existing supervised learning models ineffective at learning the…