26 papers · ranked by Valyu relevance
de Bodt, Cyril, Diaz-Papkovich, Alex + 38 more
``` Cyril de Bodt 1,2,, Alex Diaz-Papkovich 3,, Michael Bleher 4 , Kerstin Bunte 5 , Corinna Coupette 6,7,8 , Sebastian Damrich 9 , Enrique Fita Sanmartin 10,11 , Fred A. Hamprecht 12, Emoke- ˝ Agnes Horv ´ at´ 13, Dhruv Kohli 14, Smita Krishnaswamy 15 , John A. Lee 2 , Boudewijn P. F. Lelieveldt 16, Leland McInnes 17…
Srivathsan Amruth
SPREV, denoting (hyper)Sphere REduced to two-dimensional REgular Polygon for Visualisation, is a novel dimensionality reduction technique developed to addresses the challenges presented by reducing the dimension and data visualisation of labelled datasets characterized by the convergence of trifecta of…
Alexis Payton, Kyle R. Roell, Meghan E. Rebuli, William Valdar + 2 more
'Ilona Jaspers' 'Julia E. Rager'] Toxicology research has rapidly evolved, leveraging increasingly advanced technologies in high-throughput approaches to yield important information on toxicological mechanisms and health outcomes. Data produced through toxicology studies are consequently becoming larger, often…
Claire Simpson, Evgeniy Tabatsky, Zainab Rahil, Devon J. Eddins + 11 more
Unsupervised clustering is a powerful machine-learning technique widely used to analyze high-dimensional biological data. It plays a crucial role in uncovering patterns, structure, and inherent relationships within complex datasets without relying on predefined labels. In the context of biology, high-dimensional data…
Serena Hughes, Timothy Hamilton, Tom Kolokotrones, Eric J. Deeds
Manifold learning builds on the “manifold hypothesis,” which posits that data in high-dimensional datasets are drawn from lower-dimensional manifolds. Current tools generate global embeddings of data, rather than the local maps used to define manifolds mathematically. These tools also cannot assess whether the manifold…
Vamsi Manthena, Diego Jarquín, Rajeev K. Varshney, Manish Roorkiwal + 3 more
'Girish Prasad Dixit' 'Chellapilla Bharadwaj' 'Reka Howard'] The development of genomic selection (GS) methods has allowed plant breeding programs to select favorable lines using genomic data before performing field trials. Improvements in genotyping technology have yielded high-dimensional genomic marker data which…
Somya Sharma, Marten Thompson, Debra Laefer, Michael Lawler + 7 more
'Kevin McIlhany' 'Olivier Pauluis' 'Dallas R. Trinkle' 'Snigdhansu Chatterjee' 'Donald J. Jacobs' 'Emmanouil Varouchakis' 'Dionissios T. Hristopulos'] We present an overview of four challenging research areas in multiscale physics and engineering as well as four data science topics that may be developed for addressing…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Parisa Hajibabaee, Farhad Pourkamali‐Anaraki, Mohammad Amin Hariri‐Ardebili
'Mohammad Amin Hariri‐Ardebili'] Abstract—A fundamental task in machine learning involves visualizing high-dimensional data sets that arise in high-impact application domains. When considering the context of large imbalanced data, this problem becomes much more challenging. In this paper, the t-Distributed Stochastic…
Chun Kit Jeffery Hou, Kamran Behdinan
Surrogate modeling has been popularized as an alternative to full-scale models in complex engineering processes such as manufacturing and computer-assisted engineering. The modeling demand exponentially increases with complexity and number of system parameters, which consequently requires higher-dimensional engineering…
María Martínez-García, Pablo M. Olmos
The advent of high-throughput technologies has produced an increase in the dimensionality of omics datasets, which limits the application of machine learning methods due to the great unbalance between the number of observations and features. In this scenario, dimensionality reduction is essential to extract the…
Sönke Beier, Paula Pirker-Díaz, Friedrich Pagenkopf, Karoline Wiesner
Diffusion Map is a spectral dimensionality reduction technique which is able to uncover nonlinear submanifolds in high-dimensional data. And, it is increasingly applied across a wide range of scientific disciplines, such as biology, engineering, and social sciences. But data preprocessing, parameter settings and…
Hee Cheol Chung, Yang Ni, Irina Gaynanova
Sequencing-based technologies provide an abundance of high-dimensional biological data sets with highly skewed and zero-inflated measurements. Despite the computational efficiency and high interpretability offered by linear classification methods, the violation of underlying distribution assumptions, driven by high…
Marco Del Giudice, Praveen Kumar Donta
This paper introduces relative density clouds, a simple but powerful method to visualize the relative density of two groups in multivariate space. Relative density clouds employ k-nearest neighbor density estimates to provide information about group differences throughout the entire distribution of the variables. The…
Theodoulos Rodosthenous, Vahid Shahrezaei, Marina Evangelou, Shibiao Wan
'Shibiao Wan'] Non-linear dimensionality reduction can be performed by manifold learning approaches, such as stochastic neighbour embedding (SNE), locally linear embedding (LLE) and isometric feature mapping (ISOMAP). These methods aim to produce two or three latent embeddings, primarily to visualise the data in…
Chitrita Goswami, Debarka Sengupta
We introduce InGene, the first of its kind, fast and scalable non-linear, unsupervised method for analyzing single-cell RNA sequencing data (scRNA-seq). While non-linear dimensionality reduction techniques such as tSNE and UMAP are effective at visualizing cellular sub-populations in low-dimensional space, they do not…
Junning Feng, Yong Liang, Tianwei Yu
Dimension reduction is ubiquitous in high dimensional data analysis. Divergent data characteristics have driven the development of various techniques in this field. Although individual techniques can capture specific aspects of data, they often struggle to grasp all the intricate and complex patterns and structures. To…
Authors not listed
We present a gridless framework for computing high-dimensional conformational free energy surfaces (FES) of flexible molecules using enhanced sampling trajectories. By combining concurrent well-tempered metadynamics with Density Peaks Advanced (DPA) clustering, our approach bypasses the dimensionality limitations of…
Luigi Caputi, Anna Pidnebesna, Jaroslav Hlinka
This paper extends the possibility to examine the underlying curvature of data through the lens of topology by using the Betti curves, tools of Persistent Homology. We show that low-dimensional Betti curve approximations effectively distinguish not only Euclidean, but also spherical and hyperbolic geometric matrices…
Authors not listed
High-throughput experimentation (HTE) in materials science generates vast, high-dimensional datasets relating synthesis parameters to material properties. While machine learning (ML) models excel at predicting properties from these parameters, they often fail to distinguish causal drivers from merely correlated…
Guerard, Guillaume, Djebali, Sonia
The advent of the big data paradigm has revolutionized the way industries handle and analyze information, ushering in an era characterized by unprecedented volumes, velocities, and varieties of data. In this context, mixed data clustering emerges as a critical challenge, necessitating innovative approaches to…
Authors not listed
Deciphering the correct mechanism governing certain phenomenon in polyelectrolyte (PE) brush grafted systems, revealed through atomistic simulations, is an extremely challenging problem. In a recent study, our all-atom molecular dynamics (MD) simulations revealed a non-linearly large electroosmotic flow (in the…
Tendai Mapungwana Chikake, Boris Goldengorin
We introduce usage of a reduction property of penalty-based formulation of pseudo-Boolean polynomials as a mechanism for invariant dimensionality reduction in cluster analysis processes. In our experiments, we show that multidimensional data, like 4-dimensional Iris Flower dataset can be reduced to 2-dimensional space…
Mingchen Yao, Anoop Praturu, Tatyana Sharpee
The increasing size of datasets poses challenges for their visualization and interpretation, highlighting the need for scalable and effective analysis methods. Hyperbolic embedding have shown strong potential in capturing complex hierarchical structures across diverse systems. However, existing hyperbolic embedding…
Authors not listed
Quantitative Structure-Activity Relationship (QSAR) modeling is a pillar of computational drug discovery. However, standard machine learning (ML) models are often confounded by the high-dimensional and intensely correlated nature of molecular descriptors. A model may identify a "bulk" property (e.g., molecular weight)…
Authors not listed
Real-world datasets in chemical engineering and bioengineering processes--such as those from catalytic reactors, multiphase flows, polymerization reactors, bioreactors, and clinical trials--can often be unlabelled or disorganized, rendering the training of existing supervised learning models ineffective at learning the…