25 papers · ranked by Valyu relevance
Srivathsan Amruth
SPREV, denoting (hyper)Sphere REduced to two-dimensional REgular Polygon for Visualisation, is a novel dimensionality reduction technique developed to addresses the challenges presented by reducing the dimension and data visualisation of labelled datasets characterized by the convergence of trifecta of…
Alexander N. Gorban, Valery A. Makarov, Ivan Y. Tyukin
High-dimensional data and high-dimensional representations of reality are inherent features of modern Artificial Intelligence systems and applications of machine learning. The well-known phenomenon of the “curse of dimensionality” states: many problems become exponentially difficult in high dimensions. Recently, the…
Alexander N. Gorban, Valeri A. Makarov, Ivan Tyukin
- 1 Department of Mathematics, University of Leicester, Leicester LE1 7RH, UK; I.Tyukin@le.ac.uk - 2 Laboratory of Advanced Methods for High-Dimensional Data Analysis, Lobachevsky University, 603022 Nizhny Novgorod, Russia - 3 Instituto de Matemática Interdisciplinar, Faculty of Mathematics, Universidad Complutense de…
de Bodt, Cyril, Diaz-Papkovich, Alex + 38 more
``` Cyril de Bodt 1,2,, Alex Diaz-Papkovich 3,, Michael Bleher 4 , Kerstin Bunte 5 , Corinna Coupette 6,7,8 , Sebastian Damrich 9 , Enrique Fita Sanmartin 10,11 , Fred A. Hamprecht 12, Emoke- ˝ Agnes Horv ´ at´ 13, Dhruv Kohli 14, Smita Krishnaswamy 15 , John A. Lee 2 , Boudewijn P. F. Lelieveldt 16, Leland McInnes 17…
Serena Hughes, Timothy Hamilton, Tom Kolokotrones, Eric J. Deeds
Manifold learning builds on the “manifold hypothesis,” which posits that data in high-dimensional datasets are drawn from lower-dimensional manifolds. Current tools generate global embeddings of data, rather than the local maps used to define manifolds mathematically. These tools also cannot assess whether the manifold…
Vamsi Manthena, Diego Jarquín, Rajeev K. Varshney, Manish Roorkiwal + 3 more
'Girish Prasad Dixit' 'Chellapilla Bharadwaj' 'Reka Howard'] The development of genomic selection (GS) methods has allowed plant breeding programs to select favorable lines using genomic data before performing field trials. Improvements in genotyping technology have yielded high-dimensional genomic marker data which…
Shamus M. Cooley, Timothy Hamilton, J. Christian J. Ray, Eric J. Deeds
High-dimensional data are becoming increasingly common in nearly all areas of science. Developing approaches to analyze these data and understand their meaning is a pressing issue. This is particularly true for the rapidly growing field of single-cell RNA-Seq (scRNA-Seq), a technique that simultaneously measures the…
Alex Dexter, Spencer A. Thomas, Rory T. Steven, Kenneth N. Robinson + 18 more
High dimensionality omics and hyperspectral imaging datasets present difficult challenges for feature extraction and data mining due to huge numbers of features that cannot be simultaneously examined. The sample numbers and variables of these methods are constantly growing as new technologies are developed, and…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Evgeny M. Mirkes, Jeza Allohibi, Alexander Gorban
The curse of dimensionality causes the well-known and widely discussed problems for machine learning methods. There is a hypothesis that using the Manhattan distance and even fractional $l_{p}$ quasinorms (for p less than 1) can help to overcome the curse of dimensionality in classification problems. In this study, we…
Scotland C. Leman, Leanna House, Dipayan Maiti, Alex Endert + 2 more
'Chris North' 'Fabio Rapallo'] Typical data visualizations result from linear pipelines that start by characterizing data using a model or algorithm to reduce the dimension and summarize structure, and end by displaying the data in a reduced dimensional form. Sensemaking may take place at the end of the pipeline when…
María Martínez-García, Pablo M. Olmos
The advent of high-throughput technologies has produced an increase in the dimensionality of omics datasets, which limits the application of machine learning methods due to the great unbalance between the number of observations and features. In this scenario, dimensionality reduction is essential to extract the…
Chun Kit Jeffery Hou, Kamran Behdinan
Surrogate modeling has been popularized as an alternative to full-scale models in complex engineering processes such as manufacturing and computer-assisted engineering. The modeling demand exponentially increases with complexity and number of system parameters, which consequently requires higher-dimensional engineering…
Hee Cheol Chung, Yang Ni, Irina Gaynanova
Sequencing-based technologies provide an abundance of high-dimensional biological data sets with highly skewed and zero-inflated measurements. Despite the computational efficiency and high interpretability offered by linear classification methods, the violation of underlying distribution assumptions, driven by high…
Marco Del Giudice, Praveen Kumar Donta
This paper introduces relative density clouds, a simple but powerful method to visualize the relative density of two groups in multivariate space. Relative density clouds employ k-nearest neighbor density estimates to provide information about group differences throughout the entire distribution of the variables. The…
Karaj Khosla, Indra Prakash Jha, Vibhor Kumar
Dimension reduction is often used for several procedures of analysis of high dimensional biomedical data-sets such as classification or outlier detection. To improve performance of such data-mining steps, preserving both distance information and local topology among data-points could be more useful than giving priority…
Henry Kvinge, Elin Farnell, M. Kirby, Chris L. Peterson
—Dimensionality-reduction techniques are a fundamental tool for extracting useful information from highdimensional data sets. Because secant sets encode manifold geometry, they are a useful tool for designing meaningful datareduction algorithms. In one such approach, the goal is to construct a projection that maximally…
Authors not listed
We present a gridless framework for computing high-dimensional conformational free energy surfaces (FES) of flexible molecules using enhanced sampling trajectories. By combining concurrent well-tempered metadynamics with Density Peaks Advanced (DPA) clustering, our approach bypasses the dimensionality limitations of…
Suchismita Das, Nikhil R. Pal
—Here, we propose an unsupervised fuzzy rule-based dimensionality reduction method primarily for data visualization. It considers the following important issues relevant to dimensionality reduction-based data visualization: (i) preservation of neighborhood relationships, (ii) handling data on a non-linear manifold…
Authors not listed
High-throughput experimentation (HTE) in materials science generates vast, high-dimensional datasets relating synthesis parameters to material properties. While machine learning (ML) models excel at predicting properties from these parameters, they often fail to distinguish causal drivers from merely correlated…
Authors not listed
Deciphering the correct mechanism governing certain phenomenon in polyelectrolyte (PE) brush grafted systems, revealed through atomistic simulations, is an extremely challenging problem. In a recent study, our all-atom molecular dynamics (MD) simulations revealed a non-linearly large electroosmotic flow (in the…
Denis Polunin, Irina Shtaiger, Vadim Efimov
Biologists more and more have to deal with objects with non-numeric descriptions: texts (e.g. genetic sequences or even whole genomes), graphs, images, etc. There even could be no variables or descriptions at all when variability of objects is defined by similarity matrix. It is also possible to have too many variables…
J. T. Fry, Matt Slifko, Scotland Leman
Dimension reduction and visualization is a staple of data analytics. Methods such as Principal Component Analysis (PCA) and Multidimensional Scaling (MDS) provide low dimensional (LD) projections of high dimensional (HD) data while preserving an HD relationship between observations. Traditional biplots assign meaning…
Authors not listed
Quantitative Structure-Activity Relationship (QSAR) modeling is a pillar of computational drug discovery. However, standard machine learning (ML) models are often confounded by the high-dimensional and intensely correlated nature of molecular descriptors. A model may identify a "bulk" property (e.g., molecular weight)…
Authors not listed
Real-world datasets in chemical engineering and bioengineering processes--such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials--can often be unlabelled or disorganized, rendering the training of existing supervised learning models ineffective at learning the…