27 papers · ranked by Valyu relevance
Quan Wang
Principal component analysis (PCA) is a popular tool for linear dimensionality reduction and feature extraction. Kernel PCA is the nonlinear form of PCA, which better exploits the complicated spatial structure of high-dimensional features. In this paper, we first review the basic ideas of PCA and kernel PCA. Then we…
Mitja Briscik, Marie-Agnès Dillies, Sébastien Déjean
Background Kernel methods have been proven to be a powerful tool for the integration and analysis of high-throughput technologies generated data. Kernels offer a nonlinear version of any linear algorithm solely based on dot products. The kernelized version of principal component analysis is a valid nonlinear…
Miguel Alfonso Mendez
Dimensionality reduction is the essence of many data processing problems, including filtering, data compression, reduced-order modeling and pattern analysis. While traditionally tackled using linear tools in the fluid dynamics community, nonlinear tools from machine learning are becoming increasingly popular. This…
Zahra Moghaddasi, Hamid A. Jalab, Rafidah Md Noor, Saeed Aghabozorgi
Digital image forgery is becoming easier to perform because of the rapid development of various manipulation tools. Image splicing is one of the most prevalent techniques. Digital images had lost their trustability, and researches have exerted considerable effort to regain such trustability by focusing mostly on…
Wen Bo Liu, Sheng Nan Liang, Xi Wen Qin, Seyedali Mirjalili
Gene expression data has the characteristics of high dimensionality and a small sample size and contains a large number of redundant genes unrelated to a disease. The direct application of machine learning to classify this type of data will not only incur a great time cost but will also sometimes fail to improved…
Alberto García‐González, Antonio Huerta, Sergio Zlotnik, Pedro Miguel Bravo Díez
'Pedro Miguel Bravo Díez'] Methodologies for multidimensionality reduction aim at discovering low-dimensional manifolds where data ranges. Principal Component Analysis (PCA) is very effective if data have linear structure. But fails in identifying a possible dimensionality reduction if data belong to a nonlinear…
Benyamin Ghojogh, Mark Crowley
This is a detailed tutorial paper which explains the Principal Component Analysis (PCA), Supervised PCA (SPCA), kernel PCA, and kernel SPCA. We start with projection, PCA with eigendecomposition, PCA with one and multiple projection directions, properties of the projection matrix, reconstruction error minimization, and…
Christina Leitner, Franz Pernkopf
In this paper, we apply kernel PCA for speech enhancement and derive pre-image iterations for speech enhancement. Both methods make use of a Gaussian kernel. The kernel variance serves as tuning parameter that has to be adapted according to the SNR and the desired degree of de-noising. We develop a method to derive a…
Daniel Gedon, Antôni H. Ribeiro, Niklas Wahlström, Thomas B. Schön
—Kernel principal component analysis (kPCA) is a widely studied method to construct a low-dimensional data representation after a nonlinear transformation. The prevailing method to reconstruct the original input signal from kPCA—an important task for denoising—requires us to solve a supervised learning problem. In this…
Xu Liu, Yuchao Zhang, Hua Yang, Lisheng Wang + 1 more
Kernel methods, such as kernel PCA, kernel PLS, and support vector machines, are widely known machine learning techniques in biology, medicine, chemistry, and material science. Based on nonlinear mapping and Coulomb function, two 3D kernel approaches were improved and applied to predictions of the four protein tertiary…
Jérôme Mariette, Nathalie Villa-Vialaneix
Recent high-throughput sequencing advances have expanded the breadth of available omics datasets and the integrated analysis of multiple datasets obtained on the same samples has allowed to gain important insights in a wide range of applications. However, the integration of various sources of information remains a…
Ferran Reverter, Esteban Vegas, Josep M Oller
Background Nowadays, combining the different sources of information to improve the biological knowledge available is a challenge in bioinformatics. One of the most powerful methods for integrating heterogeneous data types are kernel-based methods. Kernel-based data integration approaches consist of two basic steps…
Yichuan Deng, Zhao Song, Zifan Wang, Han Zhang
Principal Component Analysis (PCA) is a widely used technique in machine learning, data analysis and signal processing. With the increase in the size and complexity of datasets, it has become important to develop low-space usage algorithms for PCA. Streaming PCA has gained significant attention in recent years, as it…
Kai Shen, Anya M. McGuirk, Yuwei Liao, Arin Chaudhuri + 1 more
'Deovrat Kakde'] Sensor data analysis plays a key role in health assessment of critical equipment. Such data are multivariate and exhibit nonlinear relationships. This paper describes how one can exploit nonlinear dimension reduction techniques, such as the t-distributed stochastic neighbor embedding (t-SNE) and kernel…
Yannis Pantazis, Christos Tselas, Kleanthi Lakiotaki, Vincenzo Lagani + 1 more
High-throughput technologies such as microarrays and RNA-sequencing (RNA-seq) allow to precisely quantify transcriptomic profiles, generating datasets that are inevitably high-dimensional. In this work, we investigate whether the whole human transcriptome can be represented in a compressed, low dimensional latent space…
Mitja Briscik, Marie‐Agnès Dillies, Sébastien Dejean
Kernel methods have been proven to be a powerful tool for the integration and analysis of highthroughput technologies generated data. Kernels offer a nonlinear version of any linear algorithm solely based on dot products. The kernelized version of Principal Component Analysis is a valid nonlinear alternative to tackle…
Hyunwook Koh
In high-dimensional omics studies, researchers often conduct kernel association testing to power-fully detect the relationship of the genetic or microbial composition with human health or disease. Especially, in human microbiome studies, its dimension reduction analysis follows to visually represent complex microbiome…
Authors not listed
Metastable states and the conformational transitions in between them are key to understanding dynamical behaviour and function of large-scale molecular systems. By combining basic dimensionality reduction techniques with a state-of-the art approximation of the Koopman operator associated to molecular dynamics…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Authors not listed
We adapted an existing approach to identifying stabilisable crystal structures from prediction sets - the Generalised Convex Hull (GCH) - to improve its application to molecular crystal structures. This was achieved by modifying the Smooth Overlap of Atomic Positions (SOAP) kernel to define the similarity of molecular…
Maria Carilli, Kayla Jackson, Lior Pachter
Contrastive learning methods can be powerful tools for genomics, enabling the identification of signals in an experiment via dimension reduction while reducing noise using a control. One such popular approach is contrastive PCA, which, despite being used in a variety of settings, does not scale to large datasets. We…
Ping Yang, E. Adrian Henle, Cory M. Simon, Xiaoli Fern
Pesticides benefit agriculture by increasing crop yield, quality, and security. However, pesticides may inadvertently harm bees, which are valuable as pollinators. Thus, candidate pesticides in development pipelines must be assessed for toxicity to bees. Leveraging a data set of 382 molecules with toxicity labels from…
Martin Seifrid, Stanley Lo, Dylan Choi, Gary Tom + 12 more
Martin Seifrid 1 , Stanley Lo 2 , Dylan G. Choi 3 , Gary Tom 2 , My Linh Le 3 , Kunyu Li 3 , Rahul Sankar 3 , Hoai-Thanh Vuong 3 , Hiba Wakidi 3 , Ahra Yi 3 , Ziyue Zhu 3 , Nora Schopp 3 , Aaron Peng 3 , Benjamin Luginbuhl 3 , Thuc-Quyen Nguyen 3 , Alán Aspuru-Guzik 2
Suk-Heung Song, Herb Ryan, Jens Hoefflin, Taeyoon Kyung + 3 more
Analytical technologies for engineered biological systems hold great promise in addressing various challenges in modern pharmaceuticals and biomedical therapies. These endeavors often follow a design-build-test-learn approach, utilizing biological data from genetic circuits, signal pathways, metabolites, and proteins…
Radu A. Talmazan, Jakob Gamper, Ivan Castillo, Thomas S. Hofer + 1 more
Supramolecular transition metal catalysts with tailored reaction environments allow for the usage of abundant 3d metals as catalytic centres, leading to more sustainable chemical processes. However, such catalysts are large and flexible systems with intricate interactions, resulting in complex reaction coordinates. To…
Philippe Boileau, Nima S. Hejazi, Sandrine Dudoit
Statistical analyses of high-throughput sequencing data have re-shaped the biological sciences. In spite of myriad advances, recovering interpretable biological signal from data corrupted by technical noise remains a prevalent open problem. Several classes of procedures, among them classical dimensionality reduction…
Oliver Christensen, Alexander Bagger, Jan Rossmeisl
For the electrochemical CO2 reduction reaction, different metal catalysts produce different products preferentially. However, the differences between the metals' reaction pathways that lead to these different products is still not fully understood. In this work, we analyze CO vs. HCOOH formation from CO2 using…