24 papers · ranked by Valyu relevance
Javier Perera-Lago, Victor Toscano-Duran, Eduardo Paluzo-Hidalgo, Rocio Gonzalez-Diaz + 2 more
In recent years, deep learning has gained popularity for its ability to solve complex classification tasks. It provides increasingly better results thanks to the development of more accurate models, the availability of huge volumes of data and the improved computational capabilities of modern computers. However, these…
He, Taotao, Luo, Jun + 2 more
Selecting an optimal subset of features or instances under an information-theoretic criterion has become an effective preprocessing strategy for reducing data complexity while preserving essential information. This study investigates two representative problems within this paradigm: feature selection based on the…
Rajat Saini, Anoop Kumar Tiwari, Abhigyan Nath, Phool Singh + 2 more
'S. P. Maurya' 'Mohd Asif Shah'] The dimension and size of data is growing rapidly with the extensive applications of computer science and lab based engineering in daily life. Due to availability of vagueness, later uncertainty, redundancy, irrelevancy, and noise, which imposes concerns in building effective learning…
Víctor Toscano-Durán, Javier Perera-Lago, Eduardo Paluzo-Hidalgo, Rocı́o González-Dı́az + 2 more
Learning Authors: ['Víctor Toscano-Durán' 'Javier Perera-Lago' 'Eduardo Paluzo-Hidalgo' 'Rocı́o González-Dı́az' 'Miguel Á. Gutiérrez-Naranjo' 'Matteo Rucco'] In recent years, Deep Learning has gained popularity for its ability to solve complex classification tasks, increasingly delivering better results thanks to the…
Handuo Zhang, Jun Na, Bin Zhang, Claudia Campolo
With the development of intelligent IoT applications, vast amounts of data are generated by various volume sensors. These sensor data need to be reduced at the sensor and then reconstructed later to save bandwidth and energy. As the reduced data increase, the reconstructed data become less accurate. Usually, the…
Matthew S. Schmitt, Maciej Koch-Janusz, Michel Fruchart, Daniel S. Seara + 2 more
Model reduction is the construction of simple yet predictive descriptions of the dynamics of many-body systems in terms of a few relevant variables. A prerequisite to model reduction is the identification of these relevant variables, a task for which no general method exists. Here, we develop a systematic approach…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Paola Patricia Ariza-Colpas, Enrico Vicario, Ana Isabel Oviedo-Carrascal, Shariq Butt Aziz + 7 more
The Assisted Living Environments Research Area-AAL (Ambient Assisted Living), focuses on generating innovative technology, products, and services to assist, medical care and rehabilitation to older adults, to increase the time in which these people can live. independently, whether they suffer from neurodegenerative…
Caroline Keller, Celine Caseys, Daniel J. Kliebenstein
Data reduction methods are frequently employed in large genomics and phenomics studies to extract core patterns, reduce dimensionality, and alleviate multiple testing effects. Principal component analysis (PCA), in particular, identifies the components that capture the most variance within omics datasets. While data…
Alain de Cheveigné
This is the second part of a two-part essay on memory and its inseparable nemesis, forgetting. It looks at memory from a computational perspective in terms of function and constraints, in the rational spirit of Marr (1982) or Anderson (1989). The core question is: How to fit an infinite past into finite storage? The…
Changqing Zhao, Ling Xia Liao, Guomin Chen, Han-Chieh Chao + 2 more
'Javier Prieto' 'Mehmet Rasit Yuce'] The accurate and efficient classification of network traffic, including malicious traffic, is essential for effective network management, cybersecurity, and resource optimization. However, traffic classification methods in modern, complex, and dynamic networks face significant…
Maximilian Woollard, Pratibha Panwar, Luke T.G. Harland, Jeremy Mo + 5 more
With the dramatic take-up of spatially resolved transcriptomics biotechnologies, performing spatially-aware analysis of the resulting data is crucial to maximise advances in biological understanding. Dimensionality reduction is a first step in almost any analysis of spatial transcriptomics data, regardless of whether…
Peng Yu, Yifeng Zheng, Ziwen Liu, Baoya Wei + 6 more
'Ziqiong Lin' 'Zhehan Li' 'Éloi Bossé' 'Sotiris Kotsiantis' 'Yong Deng'] With the development of intelligent technology, data in practical applications show exponential growth in quantity and scale. Extracting the most distinguished attributes from complex datasets becomes a crucial problem. The existing attribute…
Aidan J. Hughes, Keith Worden, Nikolaos Dervilis, Timothy J. Rogers
technologies Authors: ['Aidan J. Hughes' 'Keith Worden' 'Nikolaos Dervilis' 'Timothy J. Rogers'] Classification models are a key component of structural digital twin technologies used for supporting asset management decision-making. An important consideration when developing classification models is the dimensionality…
Humayra Tasnim, Soumya Dutta, Melanie Moses
- The research introduces a novel and adaptable method for interpreting informative features of large scale spatiotemporal data, applicable to diverse datasets from different domains. - The proposed technique identifies key informative timesteps and uses information-based fusion to summarize salient patterns of…
Tümay Capraz, Wolfgang Huber
A fundamental step in many analyses of high-dimensional data is dimension reduction. Two basic approaches are introduction of new, synthetic coordinates, and selection of extant features. Advantages of the latter include interpretability, simplicity, transferability and modularity. A common criterion for unsupervised…
Keisuke Ozawa
Statistically weighted principal component analysis (wPCA) is widely used to reduce the noise of scanning transmission electron microscopy-energy-dispersive X-ray (STEM-EDX) spectroscopy data. It is beneficial to retain the spatial resolution of observation in each step of the analysis, but the direct application of…
Tianyi Chen, Zhi‐Qin John Xu
Neural networks have been extensively applied to a variety of tasks, achieving astounding results. Applying neural networks in the scientific field is an important research direction that is gaining increasing attention. In scientific applications, the scale of neural networks is generally moderate size, mainly to…
Roberta Coletti, J. Orestes Cerdeira, Marcos Raydan, Marta B. Lopes
High-dimensional omics data often contain more variables than observations, which negatively impacts the performance of classical data analysis methods. Dimensionality reduction is typically addressed through variable selection strategies that incorporate a penalty term into the model. While effective for selecting…
Bartłomiej Fliszkiewicz, Marcin Sajdak
The aim of the following research is to assess the applicability of calculated quantum properties of molecular fragments as molecular descriptors in machine learning classification task. The research is based on bio-concentration and QM9-extended databases. A number of compounds with results from quantum-chemical…
Esther Heid, Charles J. McGill, Florence H. Vermeire, William H. Green
Characterizing uncertainty in machine learning models has recently gained interest in the context of machine learning reliability, robustness, safety, and active learning. Here, we separate the total uncertainty into contributions from noise in the data (aleatoric) and shortcomings of the model (epistemic), further…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Authors not listed
The recent release of Meta's Open Molecules 2025 dataset (OMol25) has enabled the creation of pretrained NNPs that can predict the energy of unseen molecules in a variety of charge and spin states. However, these models do not explicitly consider charge- or spin-based physics, potentially impacting the accuracy of…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…