26 papers · ranked by Valyu relevance
Srivathsan Amruth
SPREV, denoting (hyper)Sphere REduced to two-dimensional REgular Polygon for Visualisation, is a novel dimensionality reduction technique developed to addresses the challenges presented by reducing the dimension and data visualisation of labelled datasets characterized by the convergence of trifecta of…
Alexander N. Gorban, Valery A. Makarov, Ivan Y. Tyukin
High-dimensional data and high-dimensional representations of reality are inherent features of modern Artificial Intelligence systems and applications of machine learning. The well-known phenomenon of the “curse of dimensionality” states: many problems become exponentially difficult in high dimensions. Recently, the…
Alexander N. Gorban, Valeri A. Makarov, Ivan Tyukin
- 1 Department of Mathematics, University of Leicester, Leicester LE1 7RH, UK; I.Tyukin@le.ac.uk - 2 Laboratory of Advanced Methods for High-Dimensional Data Analysis, Lobachevsky University, 603022 Nizhny Novgorod, Russia - 3 Instituto de Matemática Interdisciplinar, Faculty of Mathematics, Universidad Complutense de…
de Bodt, Cyril, Diaz-Papkovich, Alex + 38 more
``` Cyril de Bodt 1,2,, Alex Diaz-Papkovich 3,, Michael Bleher 4 , Kerstin Bunte 5 , Corinna Coupette 6,7,8 , Sebastian Damrich 9 , Enrique Fita Sanmartin 10,11 , Fred A. Hamprecht 12, Emoke- ˝ Agnes Horv ´ at´ 13, Dhruv Kohli 14, Smita Krishnaswamy 15 , John A. Lee 2 , Boudewijn P. F. Lelieveldt 16, Leland McInnes 17…
Jacob M. Graving, Iain D. Couzin
Scientific datasets are growing rapidly in scale and complexity. Consequently, the task of understanding these data to answer scientific questions increasingly requires the use of compression algorithms that reduce dimensionality by combining correlated features and cluster similar observations to summarize large…
Vamsi Manthena, Diego Jarquín, Rajeev K. Varshney, Manish Roorkiwal + 3 more
'Girish Prasad Dixit' 'Chellapilla Bharadwaj' 'Reka Howard'] The development of genomic selection (GS) methods has allowed plant breeding programs to select favorable lines using genomic data before performing field trials. Improvements in genotyping technology have yielded high-dimensional genomic marker data which…
Shamus M. Cooley, Timothy Hamilton, J. Christian J. Ray, Eric J. Deeds
High-dimensional data are becoming increasingly common in nearly all areas of science. Developing approaches to analyze these data and understand their meaning is a pressing issue. This is particularly true for the rapidly growing field of single-cell RNA-Seq (scRNA-Seq), a technique that simultaneously measures the…
Somya Sharma, Marten Thompson, Debra Laefer, Michael Lawler + 7 more
'Kevin McIlhany' 'Olivier Pauluis' 'Dallas R. Trinkle' 'Snigdhansu Chatterjee' 'Donald J. Jacobs' 'Emmanouil Varouchakis' 'Dionissios T. Hristopulos'] We present an overview of four challenging research areas in multiscale physics and engineering as well as four data science topics that may be developed for addressing…
Alex Dexter, Spencer A. Thomas, Rory T. Steven, Kenneth N. Robinson + 18 more
High dimensionality omics and hyperspectral imaging datasets present difficult challenges for feature extraction and data mining due to huge numbers of features that cannot be simultaneously examined. The sample numbers and variables of these methods are constantly growing as new technologies are developed, and…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Chun Kit Jeffery Hou, Kamran Behdinan
Surrogate modeling has been popularized as an alternative to full-scale models in complex engineering processes such as manufacturing and computer-assisted engineering. The modeling demand exponentially increases with complexity and number of system parameters, which consequently requires higher-dimensional engineering…
Evgeny M. Mirkes, Jeza Allohibi, Alexander Gorban
The curse of dimensionality causes the well-known and widely discussed problems for machine learning methods. There is a hypothesis that using the Manhattan distance and even fractional $l_{p}$ quasinorms (for p less than 1) can help to overcome the curse of dimensionality in classification problems. In this study, we…
Marco Del Giudice, Praveen Kumar Donta
This paper introduces relative density clouds, a simple but powerful method to visualize the relative density of two groups in multivariate space. Relative density clouds employ k-nearest neighbor density estimates to provide information about group differences throughout the entire distribution of the variables. The…
Brian Cleary, Le Cong, Eric S. Lander, Aviv Regev
RNA profiling is an excellent phenotype of cellular responses and tissue states, but can be costly to generate at the massive scale required for studies of regulatory circuits, genetic states or perturbation screens. Here, we draw on a series of advances over the last decade in the field of mathematics to establish a…
Theodoulos Rodosthenous, Vahid Shahrezaei, Marina Evangelou, Shibiao Wan
'Shibiao Wan'] Non-linear dimensionality reduction can be performed by manifold learning approaches, such as stochastic neighbour embedding (SNE), locally linear embedding (LLE) and isometric feature mapping (ISOMAP). These methods aim to produce two or three latent embeddings, primarily to visualise the data in…
Junning Feng, Yong Liang, Tianwei Yu
Dimension reduction is ubiquitous in high dimensional data analysis. Divergent data characteristics have driven the development of various techniques in this field. Although individual techniques can capture specific aspects of data, they often struggle to grasp all the intricate and complex patterns and structures. To…
Karaj Khosla, Indra Prakash Jha, Vibhor Kumar
Dimension reduction is often used for several procedures of analysis of high dimensional biomedical data-sets such as classification or outlier detection. To improve performance of such data-mining steps, preserving both distance information and local topology among data-points could be more useful than giving priority…
Authors not listed
We present a gridless framework for computing high-dimensional conformational free energy surfaces (FES) of flexible molecules using enhanced sampling trajectories. By combining concurrent well-tempered metadynamics with Density Peaks Advanced (DPA) clustering, our approach bypasses the dimensionality limitations of…
Parisa Hajibabaee, Farhad Pourkamali‐Anaraki, Mohammad Amin Hariri‐Ardebili
'Mohammad Amin Hariri‐Ardebili'] Abstract—A fundamental task in machine learning involves visualizing high-dimensional data sets that arise in high-impact application domains. When considering the context of large imbalanced data, this problem becomes much more challenging. In this paper, the t-Distributed Stochastic…
Suchismita Das, Nikhil R. Pal
—Here, we propose an unsupervised fuzzy rule-based dimensionality reduction method primarily for data visualization. It considers the following important issues relevant to dimensionality reduction-based data visualization: (i) preservation of neighborhood relationships, (ii) handling data on a non-linear manifold…
Yannis Pantazis, Christos Tselas, Kleanthi Lakiotaki, Vincenzo Lagani + 1 more
High-throughput technologies such as microarrays and RNA-sequencing (RNA-seq) allow to precisely quantify transcriptomic profiles, generating datasets that are inevitably high-dimensional. In this work, we investigate whether the whole human transcriptome can be represented in a compressed, low dimensional latent space…
Authors not listed
High-throughput experimentation (HTE) in materials science generates vast, high-dimensional datasets relating synthesis parameters to material properties. While machine learning (ML) models excel at predicting properties from these parameters, they often fail to distinguish causal drivers from merely correlated…
Authors not listed
Deciphering the correct mechanism governing certain phenomenon in polyelectrolyte (PE) brush grafted systems, revealed through atomistic simulations, is an extremely challenging problem. In a recent study, our all-atom molecular dynamics (MD) simulations revealed a non-linearly large electroosmotic flow (in the…
Nitish Bahadur, Randy Paffenroth
—Dimension Estimation (DE) and Dimension Reduction (DR) are two closely related topics, but with quite different goals. In DE, one attempts to estimate the intrinsic dimensionality or number of latent variables in a set of measurements of a random vector. However, in DR, one attempts to project a random vector, either…
Authors not listed
Quantitative Structure-Activity Relationship (QSAR) modeling is a pillar of computational drug discovery. However, standard machine learning (ML) models are often confounded by the high-dimensional and intensely correlated nature of molecular descriptors. A model may identify a "bulk" property (e.g., molecular weight)…
Authors not listed
Real-world datasets in chemical engineering and bioengineering processes--such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials--can often be unlabelled or disorganized, rendering the training of existing supervised learning models ineffective at learning the…