28 papers · ranked by Valyu relevance
Norazman Shahar, Muhammad Amir As’ari, Mohamad Hazwan Mohd Ghazali, Nasharuddin Zainal + 8 more
Accurate recognition of complex human activities from wearable sensors plays a critical role in sports analytics and human performance monitoring. However, the high dimensionality and redundancy of raw inertial data can hinder model performance and interpretability. This study proposes a hybrid feature selection…
Hengrui Luo, Jeremy E. Purvis, Didong Li
Modern datasets often exhibit high dimensionality, yet the data reside in low-dimensional manifolds that can reveal underlying geometric structures critical for data analysis. A prime example of such a dataset is a collection of cell cycle measurements, where the inherently cyclical nature of the process can be…
Yuta Hozumi, Rui Wang, Guo-Wei Wei
Most dimensionality reduction methods employ frequency domain representations obtained from matrix diagonalization and may not be efficient for large datasets with relatively high intrinsic dimensions. To address this challenge, correlated clustering and projection (CCP) offers a novel data domain strategy that does…
Antonio Di Noia, Federico Ravenda, Antonietta Mira
Dimensionality reduction is a fundamental task in modern data science. Several projection methods specifically tailored to take into account the non-linearity of the data via local embeddings have been proposed. Such methods are often based on local neighbourhood structures and require tuning the number of neighbours…
Chunli Xiang, Jing Zhou, Wen Zhou, Yongquan Zhou
With the explosive growth of data across various fields, effective data preprocessing has become increasingly critical. Evolutionary and swarm intelligence algorithms have shown considerable potential in feature selection. However, their performance often deteriorates in large-scale problems, due to premature…
Fritz Lekschas, Nezar Abdennur
Understanding high-dimensional data requires projecting it into lower-dimensional spaces, but any single projection inevitably loses information or introduces distortions. Tours address this limitation through animation of 2D projection sequences, yet existing tools present tradeoffs in the freedom and steerability of…
Sönke Beier, Paula Pirker-Díaz, Friedrich Pagenkopf, Karoline Wiesner
Diffusion Map is a spectral dimensionality reduction technique which is able to uncover nonlinear submanifolds in high-dimensional data. And, it is increasingly applied across a wide range of scientific disciplines, such as biology, engineering, and social sciences. But data preprocessing, parameter settings and…
Hezzal Kucukselbes, Ebru Sayilgan
This study introduces a real-time processing framework for decoding motor imagery EEG signals by integrating manifold learning techniques with shallow classifiers. EEG recordings were obtained from six healthy participants performing five distinct wrist and hand motor imagery tasks. To address the challenges of high…
Jianbin Tan, Pixu Shi
Understanding associations between paired high-dimensional longitudinal datasets is a fundamental yet challenging problem that arises across scientific domains, including longitudinal multi-omic studies. The difficulty stems from the complex, timevarying cross-covariance structure coupled with high dimensionality…
Aolin Chen, Pengfei Pan, Ning Quan, Bo Zhou
Feature selection is a fundamental yet challenging task in machine learning, particularly in high-dimensional settings. Although swarm intelligence and evolutionary computation methods, including ant colony optimization and grey wolf optimizer, have shown promising performance in feature selection, they still face two…
Bingxue An, Tiffany M. Tang
A plethora of dimension reduction methods have been developed to visualize highdimensional data in low dimensions. However, different dimension reduction methods often output different and possibly conflicting visualizations of the same data. This problem is further exacerbated by the choice of hyperparameters, which…
Authors not listed
High-throughput experimentation (HTE) in materials science generates vast, high-dimensional datasets relating synthesis parameters to material properties. While machine learning (ML) models excel at predicting properties from these parameters, they often fail to distinguish causal drivers from merely correlated…
Nandini Chatterjee, Aleksandr Taraskin, Hridya Divakaran, Natalia Jaeger + 3 more
The rapid evolution of single-cell technologies has generated vast, multimodal datasets encompassing genomic, transcriptomic, proteomic, and spatial information. However, high dimensionality, noise, and computational costs pose significant challenges, often introducing bias through traditional feature selection…
Rossana O. Souza, Wellington Francisco Rodrigues, Bráulio R. G. M. Couto, Marcos A. dos Santos
Logistic regression remains a widely used classification method due to its interpretability and computational efficiency, but its direct application to high-dimensional biomedical data is limited when the number of features greatly exceeds the number of samples. In this paper, we propose a reformulated logistic…
Guerard, Guillaume, Djebali, Sonia
The advent of the big data paradigm has revolutionized the way industries handle and analyze information, ushering in an era characterized by unprecedented volumes, velocities, and varieties of data. In this context, mixed data clustering emerges as a critical challenge, necessitating innovative approaches to…
Odagiu, Patrick, Belis, Vasilis + 14 more
Data sets that are specified by a large number of features are currently outside the area of applicability for quantum machine learning algorithms. An immediate solution to this impasse is the application of dimensionality reduction methods before passing the data to the quantum algorithm. We investigate six…
Authors not listed
Quantitative Structure-Activity Relationship (QSAR) modeling is a pillar of computational drug discovery. However, standard machine learning (ML) models are often confounded by the high-dimensional and intensely correlated nature of molecular descriptors. A model may identify a "bulk" property (e.g., molecular weight)…
Jinpu Cai, Yuxuan Wang, Yunhao Qiao, Cheng Wang + 7 more
Single-cell and spatial transcriptomics provide high-resolution cellular characterization, yet standard analytical approaches remain theoretically misaligned with the probabilistic nature of the data. After UMI normalization, current pipelines rely on Euclidean or log-transformed Euclidean distance for similarity…
Nicholas Markarian, Barbara E. Engelhardt, Niles A. Pierce, Paul W. Sternberg + 1 more
Principal component analysis (PCA) and k-means clustering are two seemingly different methods for dimension reduction and clustering, respectively, but can be understood as special cases of inference in a Gaussian latent variable model framework. We leverage this insight to develop a probabilistic framework and methods…
Hyeon Jeon
Dimensionality reduction (DR) is one of the most commonly used yet most easily misinterpreted tools in visual analytics. Visual analytics using DR can thus easily be unreliable: insights derived from analysis may not accurately reflect the underlying data, potentially leading to flawed knowledge and decision-making.…
George Hutchings, Pantelis Samartsidis, Corinne Donnay, Laura Gaetano + 5 more
Probabilistic latent variable models are a powerful tool for uncovering structure in high-dimensional datasets, particularly in biomedical applications. The increasing availability of large-scale epidemiological studies, such as the UK Biobank, poses important modelling challenges, including mixed data types, high…
Maria Carilli, Kayla Jackson, Lior Pachter
Contrastive learning methods can be powerful tools for genomics, enabling the identification of signals in an experiment via dimension reduction while reducing noise using a control. One such popular approach is contrastive PCA, which, despite being used in a variety of settings, does not scale to large datasets. We…
Sina Kanannejad, Noemi Bongiorni, Elisa Nordera, Sara Redaelli + 4 more
Single-cell RNA sequencing has enabled the construction of comprehensive cell atlases, yet the quality and coherence of the cell-type annotations within these atlases remain largely unexamined. When a label is applied to a transcriptionally heterogeneous population, the downstream analyses that depend on it, and…
Authors not listed
Metastable states and the conformational transitions in between them are key to understanding dynamical behaviour and function of large-scale molecular systems. By combining basic dimensionality reduction techniques with a state-of-the art approximation of the Koopman operator associated to molecular dynamics…
Authors not listed
Early on in the emergence of virtual high-throughput screening (VHTS), it was recognized that for validation to be robust and reliable, decoys should match actives as closely as possible in as many aspects as possible. This has given rise to several generations of validation sets that address previously reported…
Authors not listed
The discovery of chemically novel or structurally anomalous metal-organic frameworks (MOFs) is essential for expanding reticular design space and enhancing dataset reliability. We present CHEM-AD (Chemically Unusual Metal–organic Frameworks via Autoencoder-based Detection), a label-free, CPU-efficient pipeline that…
Christos Chatzis, David Horner, Rasmus Bro, Ann-Marie Malby Schoos + 2 more
Temporal multivariate data is ubiquitous in many domains, for instance, being collected over time at planned visits (every few months/years) in longitudinal cohorts, or every few minutes/hours in challenge tests. The analysis of such data often focuses on revealing the underlying temporal patterns common across…
Authors not listed
We present a theoretical and computational framework for virtual mass spectrometry based on Molecular Maxwell Demons (MMDs) operating as information catalysts. Building on the biological Maxwell demon framework, we demonstrate that mass spectrometry data contain categorical state information that is fundamentally…