22 papers · ranked by Valyu relevance
Gift Nyamundanda, Lorraine Brennan, Isobel Claire Gormley
Background Data from metabolomic studies are typically complex and high-dimensional. Principal component analysis (PCA) is currently the most widely used statistical technique for analyzing metabolomic data. However, PCA is limited by the fact that it is not based on a statistical model. Results Here, probabilistic…
Esra Pamukçu, Hamparsum Bozdogan, Sinan Çalık
Gene expression data typically are large, complex, and highly noisy. Their dimension is high with several thousand genes (i.e., features) but with only a limited number of observations (i.e., samples). Although the classical principal component analysis (PCA) method is widely used as a first standard step in dimension…
Anahita Nodehi, Mousa Golalizadeh, Mehdi Maadooliat, Claudio Agostinelli
'Claudio Agostinelli'] One of the most common problems that any technique encounters is the high dimensionality of the input data. This yields several problems in the subsequently performed statistical methods due to the so-called "curse of dimensionality". Several dimension reduction methods have been proposed in the…
Aman Agrawal, Alec M. Chiu, Minh Le, Eran Halperin + 1 more
Principal component analysis (PCA) is a key tool for understanding population structure and controlling for population stratification in genome-wide association studies (GWAS). With the advent of large-scale datasets of genetic variation, there is a need for methods that can compute principal components (PCs) with…
Akash Yadav, Ruda Zhang
This paper proposes a probabilistic model of subspaces based on the probabilistic principal component analysis (PCA). Given a sample of vectors in the embedding space—commonly known as a snapshot matrix—this method uses quantities derived from the probabilistic PCA to construct distributions of the sample matrix, as…
Han-Lin Hsieh, Maryam M. Shanechi
Dimensionality reduction is critical across various domains of science including neuroscience. Probabilistic Principal Component Analysis (PPCA) is a prominent dimensionality reduction method that provides a probabilistic approach unlike the deterministic approach of PCA and serves as a connection between PCA and…
Aman Agrawal, Alec M. Chiu, Minh Le, Eran Halperin + 2 more
'Sriram Sankararaman' 'Simon Gravel'] Principal component analysis (PCA) is a key tool for understanding population structure and controlling for population stratification in genome-wide association studies (GWAS). With the advent of large-scale datasets of genetic variation, there is a need for methods that can…
Mohammad Nabhan, Yajun Mei, Jianjun Shi
High dimensional data has introduced challenges that are difficult to address when attempting to implement classical approaches of statistical process control. This has made it a topic of interest for research due in recent years. However, in many cases, data sets have underlying structures, such as in advanced…
Guihong Wan, Crystal Maung, Haim Schweitzer
—Classical Principal Component Analysis (PCA) approximates data in terms of projections on a small number of orthogonal vectors. There are simple procedures to efficiently compute various functions of the data from the PCA approximation. The most important function is arguably the Euclidean distance between data items…
Didong Li, Andrew Jones, Barbara E. Engelhardt
Dimension reduction is useful for exploratory data analysis. In many applications, it is of interest to discover variation that is enriched in a "foreground" dataset relative to a "background" dataset. Recently, contrastive principal component analysis (CPCA) was proposed for this setting. However, the lack of a formal…
Divyaratan Popli, Benjamin M. Peter
Principal component analysis (PCA) and F-statistics are routinely used in population genetic and archaeogenetic studies. Here, we present a statistical framework to combine them into a joint analysis, showing where they coincide, and where slightly different assumptions made can lead to different outcomes. In…
Nicholas Markarian, Barbara E. Engelhardt, Niles A. Pierce, Paul W. Sternberg + 1 more
Principal component analysis (PCA) and k-means clustering are two seemingly different methods for dimension reduction and clustering, respectively, but can be understood as special cases of inference in a Gaussian latent variable model framework. We leverage this insight to develop a probabilistic framework and methods…
G. Durif, L. Modolo, J. E. Mold, S. Lambert-Lacroix + 1 more
The development of high throughput single-cell technologies now allows the investigation of the genome-wide diversity of transcription. This diversity has shown two faces: the expression dynamics (gene to gene variability) can be quantified more accurately, thanks to the measurement of lowly-expressed genes. Second…
Maria Carilli, Kayla Jackson, Lior Pachter
Contrastive learning methods can be powerful tools for genomics, enabling the identification of signals in an experiment via dimension reduction while reducing noise using a control. One such popular approach is contrastive PCA, which, despite being used in a variety of settings, does not scale to large datasets. We…
Heather J. Zhou, Lei Li, Yumei Li, Wei Li + 1 more
Estimating and accounting for hidden variables is widely practiced as an important step in quantitative trait locus (QTL) analysis for improving the power of QTL identification. Here we benchmark popular hidden variable inference methods including surrogate variable analysis (SVA), probabilistic estimation of…
Jochen Görtler, Thilo Spinner, Dirk Streeb, Daniel Weiskopf + 1 more
[_page_7_Figure_13.jpeg]: Multiple comparison of our method to a sampling-based approach using the Hellinger distance for input data with two to 12 dimensions. The x axis shows the increasing number of samples that were used for the sampling strategy, while the y axis shows the distance to the result from our method.…
Maria Carilli, Kayla Jackson, Lior Pachter
Contrastive learning methods can be powerful tools for genomics, enabling the identification of signals in an experiment via dimension reduction while reducing noise using a control. One such popular approach is contrastive PCA, which, despite being used in a variety of settings, does not scale to large datasets. We…
Samuel Renaud, Rachael Mansbach
Current antibacterial treatments cannot overcome the rapidly growing resistance of bacteria to antibiotic drugs, and novel treatment methods are required. One option is the development of new antimicrobial peptides (AMPs), to which bacterial resistance build-up is comparatively slow. Deep generative models have…
Oliver Christensen, Alexander Bagger, Jan Rossmeisl
For the electrochemical CO2 reduction reaction, different metal catalysts produce different products preferentially. However, the differences between the metals' reaction pathways that lead to these different products is still not fully understood. In this work, we analyze CO vs. HCOOH formation from CO2 using…
Radu A. Talmazan, Jakob Gamper, Ivan Castillo, Thomas S. Hofer + 1 more
Supramolecular transition metal catalysts with tailored reaction environments allow for the usage of abundant 3d metals as catalytic centres, leading to more sustainable chemical processes. However, such catalysts are large and flexible systems with intricate interactions, resulting in complex reaction coordinates. To…
Keisuke Ozawa
Statistically weighted principal component analysis (wPCA) is widely used to reduce the noise of scanning transmission electron microscopy-energy-dispersive X-ray (STEM-EDX) spectroscopy data. It is beneficial to retain the spatial resolution of observation in each step of the analysis, but the direct application of…
Authors not listed
Machine learning models are increasingly applied to heterogeneous materials datasets spanning different synthesis routes, measurement protocols, and structural classes. Although multi-task and representation-learning approaches are commonly used to improve predictive performance, the latent representations learned by…