25 papers · ranked by Valyu relevance
Alfredo Nava Tudela
The weak-` p norm can be used to define a measure s of sparsity. When we compute s for the discrete cosine transform coefficients of a signal, the value of s is related to the information content of said signal. We use this value of s to define a reference-free index E, called the sparsity index, that we can use to…
Hyonho Chun, Sündüz Keleş
Partial least squares regression has been an alternative to ordinary least squares for handling multicollinearity in several areas of scientific research since the 1960s. It has recently gained much attention in the analysis of high dimensional genomic data. We show that known asymptotic consistency of the partial…
Florian Huber, Julian Pollmann
Quantifying molecular similarity is a cornerstone of cheminformatics, underpinning applications from virtual screening to chemical space visualization. A wide range of molecular fingerprints and similarity metrics, most notably Tanimoto scores, are employed, but their effectiveness is highly context-dependent. In this…
Riyasat Ohib, Nicolas Gillis, Niccolò Dalmasso, Sameena Shah + 2 more
'Vamsi K. Potluru' 'Sergey M. Plis'] As evident from deep learning, very large models bring improvements in training dynamics and representation power. Yet, smaller models have benefits of energy efficiency and interpretability. To get the benefits from both ends of the spectrum we often encourage sparsity in the…
Dimitris Papailiopoulos, Alexandros G. Dimakis, Stavros Korokythakis
We introduce a novel algorithm that computes the k-sparse principal component of a positive semidefinite matrix A. Our algorithm is combinatorial and operates by examining a discrete set of special vectors lying in a low-dimensional eigen-subspace of A. We obtain provable approximation guarantees that depend on the…
Ines Wilms, Christophe Croux
Background Canonical correlation analysis (CCA) is a multivariate statistical method which describes the associations between two sets of variables. The objective is to find linear combinations of the variables in each data set having maximal correlation. In genomics, CCA has become increasingly important to estimate…
S. Park, E. Ceulemans, K. Van Deun
Principal component analysis (PCA) is an important tool for analyzing large collections of variables. It functions both as a pre-processing tool to summarize many variables into components and as a method to reveal structure in data. Different coefficients play a central role in these two uses. One focuses on the…
Gitta Kutyniok
Compressed sensing is a novel research area, which was introduced in 2006, and since then has already become a key concept in various areas of applied mathematics, computer science, and electrical engineering. It surprisingly predicts that high-dimensional signals, which allow a sparse representation by a suitable…
Agniva Chowdhury, Aritra Bose, Samson Zhou, David P. Woodruff + 1 more
Principal component analysis (PCA) is a widely used dimensionality reduction technique in machine learning and multivariate statistics. To improve the interpretability of PCA, various approaches to obtain sparse principal direction loadings have been proposed, which are termed Sparse Principal Component Analysis…
Won-Seok Lee, Hyoung-Kyu Song, Gianmarco Romano
This paper proposes an efficient channel information feedback scheme to reduce the feedback overhead of multi-user multiple-input multiple-output (MU-MIMO) hybrid beamforming systems. As massive machine type communication (mMTC) was considered in the deployments of 5G, a transmitter of the hybrid beamforming system…
Yanbo Lian, Anthony N. Burkitt
Sparse coding, predictive coding and divisive normalization have each been found to be principles that underlie the function of neural circuits in many parts of the brain, supported by substantial experimental evidence. However, the connections between these related principles are still poorly understood. In this…
Taro Tezuka
In biological neural networks, it is widely accepted that the spikes are the fundamental building blocks of information representation [1]. In contrast, whether such building blocks exist at a higher level in terms of time and in a population of neurons is a topic of ongoing debate. One approach for finding candidates…
Aydın Buluç
Multiplication of a sparse matrix with another (dense or sparse) matrix is a fundamental operation that captures the computational patterns of many data science applications, including but not limited to graph algorithms, sparsely connected neural networks, graph neural networks, clustering, and many-to-many…
Vincent Guillemot, Derek Beaton, Arnaud Gloaguen, Tommy Löfstedt + 5 more
'Brian Levine' 'Nicolas Raymond' 'Arthur Tenenhaus' 'Hervé Abdi' 'Shyamal D Peddada'] We propose a new sparsification method for the singular value decomposition-called the constrained singular value decomposition (CSVD)-that can incorporate multiple constraints such as sparsification and orthogonality for the left and…
Laura Rebollo‐Neira, Miroslav Rozložńık, Pradip Sasmal
The convergence and numerical analysis of a low memory implementation of the Orthogonal Matching Pursuit greedy strategy, which is termed Self Projected Matching Pursuit, is presented. This approach provides an iterative way of solving the least squares problem with much less storage requirement than direct linear…
Sanjar Adilov
Machine learning models for molecular-property prediction typically work with molecular representations in the form of fingerprints, descriptors, or graphs. In case of fingerprints and descriptors, molecular representations usually comprise thousands of features, which causes the curse of dimensionality for many…
Elizabeth Herbert, Srdjan Ostojic
Neural population dynamics are often highly coordinated, allowing task-related computations to be understood as neural trajectories through low-dimensional subspaces. How the network connectivity and input structure give rise to such activity can be investigated with the aid of low-rank recurrent neural networks, a…
Mina Ghashami, Edo Liberty, Jeff M. Phillips
This paper describes Sparse Frequent Directions, a variant of Frequent Directions for sketching sparse matrices. It resembles the original algorithm in many ways: both receive the rows of an input matrix An×d one by one in the streaming setting and compute a small sketch B ∈ R `×d . Both share the same strong (provably…
Tagir Akhmetshin, Arkadii Lin, Timur Madzhidov, Alexandre Varnek
Autoencoders represent a promising technique for the inverse quantitative structure-activity relationship (QSAR) task. However, undesirable bias, such as atom ordering, affects the neighbourhood behaviour of autoencoders’ latent space and, consequently, usage of the latent vectors as variables in machine-learning…
Ping Yang, E. Adrian Henle, Cory M. Simon, Xiaoli Fern
Pesticides benefit agriculture by increasing crop yield, quality, and security. However, pesticides may inadvertently harm bees, which are valuable as pollinators. Thus, candidate pesticides in development pipelines must be assessed for toxicity to bees. Leveraging a data set of 382 molecules with toxicity labels from…
Kelsey Hatzell, Yanjie Zheng
X-ray Computed Tomography (CT) is a non-invasive, non-destructive approach to imaging materials, material systems and engineered components in two- and three- dimensions. Acquisition of 3D images requires the collection of hundreds or thousands of through-thickness X-ray radiographic images from different angles. Such…
Chen Qu, Paul Houston, Qi Yu, Riccardo Conte + 3 more
Hamiltonian matrices in electronic and nuclear contexts are highly compute-intensive to calculate, mainly due to the cost for the potential matrix. Typically these matrices contain many off-diagonal elements that are orders of magnitude smaller than diagonal elements. We illustrate that here for vibrational H-matrices…
Maxat Kulmanov, Senay Kafkas, Andreas Karwath, Alexander Malic + 3 more
Recent developments in machine learning have lead to a rise of large number of methods for extracting features from structured data. The features are represented as a vectors and may encode for some semantic aspects of data. They can be used in a machine learning models for different tasks or to compute similarities…
Eric Hermes, Khachik Sargsyan, Habib Najm, Judit Zádor
We present a new algorithm for the optimization of molecular structures to saddle points on the potential energy surface using a redundant internal coordinate system. This algorithm automates the procedure of defining the internal coordinate system, including the handling of linear bending angles, e.g. through the…
Daniela Dolciami, Robert Ziolek, Daniel Davies, Michael Carter + 2 more
Chemical diversity is challenging to describe objectively. Despite this, various notions of chemical diversity are used throughout the medicinal chemistry optimization process in drug discovery. In this work, we show the usefulness of considering exploited vectors during different phases of the drug design process to…