26 papers · ranked by Valyu relevance
Javier Perera-Lago, Victor Toscano-Duran, Eduardo Paluzo-Hidalgo, Rocio Gonzalez-Diaz + 2 more
In recent years, deep learning has gained popularity for its ability to solve complex classification tasks. It provides increasingly better results thanks to the development of more accurate models, the availability of huge volumes of data and the improved computational capabilities of modern computers. However, these…
Víctor Toscano-Durán, Javier Perera-Lago, Eduardo Paluzo-Hidalgo, Rocı́o González-Dı́az + 2 more
Learning Authors: ['Víctor Toscano-Durán' 'Javier Perera-Lago' 'Eduardo Paluzo-Hidalgo' 'Rocı́o González-Dı́az' 'Miguel Á. Gutiérrez-Naranjo' 'Matteo Rucco'] In recent years, Deep Learning has gained popularity for its ability to solve complex classification tasks, increasingly delivering better results thanks to the…
Liam Steadman, Nathan Griffiths, Stephen A. Jarvis, Mark Bell + 2 more
'Shaun Helman' 'Caroline Wallbank'] Analysing and learning from spatio-temporal datasets is an important process in many domains, including transportation, healthcare and meteorology. In particular, data collected by sensors in the environment allows us to understand and model the processes acting within the…
Matthew S. Schmitt, Maciej Koch-Janusz, Michel Fruchart, Daniel S. Seara + 2 more
Model reduction is the construction of simple yet predictive descriptions of the dynamics of many-body systems in terms of a few relevant variables. A prerequisite to model reduction is the identification of these relevant variables, a task for which no general method exists. Here, we develop a systematic approach…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Caroline Keller, Celine Caseys, Daniel J. Kliebenstein
Data reduction methods are frequently employed in large genomics and phenomics studies to extract core patterns, reduce dimensionality, and alleviate multiple testing effects. Principal component analysis (PCA), in particular, identifies the components that capture the most variance within omics datasets. While data…
M. K. Alam, Azrina Abd Aziz, S. A. Latif, Azlan Awang
A wireless sensor network (WSN) deploys hundreds or thousands of nodes that may introduce large-scale data over time. Dealing with such an amount of collected data is a real challenge for energy-constraint sensor nodes. Therefore, numerous research works have been carried out to design efficient data clustering…
Changqing Zhao, Ling Xia Liao, Guomin Chen, Han-Chieh Chao + 2 more
'Javier Prieto' 'Mehmet Rasit Yuce'] The accurate and efficient classification of network traffic, including malicious traffic, is essential for effective network management, cybersecurity, and resource optimization. However, traffic classification methods in modern, complex, and dynamic networks face significant…
Canchen Li
—Data mining is about obtaining new knowledge from existing datasets. However, the data in the existing datasets can be scattered, noisy, and even incomplete. Although lots of effort is spent on developing or fine-tuning data mining models to make them more robust to the noise of the input data, their qualities still…
Anna Konstorum, Nathan Jekel, Emily Vidal, Reinhard Laubenbacher
Mass cytometry, also known as CyTOF, is a newly developed technology for quantification and classification of immune cells that can allow for analysis of over three dozen protein markers per cell. The high dimensional data that is generated requires innovative methods for analysis and visualization. We conducted a…
Karaj Khosla, Indra Prakash Jha, Vibhor Kumar
Dimension reduction is often used for several procedures of analysis of high dimensional biomedical data-sets such as classification or outlier detection. To improve performance of such data-mining steps, preserving both distance information and local topology among data-points could be more useful than giving priority…
Yasset Perez-Riverol, Max Kun, Juan Antonio Vizcaíno, Marc-Phillip Hitz + 1 more
We are moving into the age of ‘Big Data’ in biomedical research and bioinformatics. This trend could be encapsulated in this simple formula: D = S × F, where the volume of data generated (D) increases in both dimensions: the number of samples (S) and the number of sample features (F). Frequently, a typical…
Sean H. Merritt, Alexander P. Christensen
Developing interpretable machine learning models has become an increasingly important issue. One way in which data scientists have been able to develop interpretable models has been to use dimension reduction techniques. In this paper, we examine several dimension reduction techniques including two recent approaches…
Peng Yu, Yifeng Zheng, Ziwen Liu, Baoya Wei + 6 more
'Ziqiong Lin' 'Zhehan Li' 'Éloi Bossé' 'Sotiris Kotsiantis' 'Yong Deng'] With the development of intelligent technology, data in practical applications show exponential growth in quantity and scale. Extracting the most distinguished attributes from complex datasets becomes a crucial problem. The existing attribute…
Aidan J. Hughes, Keith Worden, Nikolaos Dervilis, Timothy J. Rogers
technologies Authors: ['Aidan J. Hughes' 'Keith Worden' 'Nikolaos Dervilis' 'Timothy J. Rogers'] Classification models are a key component of structural digital twin technologies used for supporting asset management decision-making. An important consideration when developing classification models is the dimensionality…
Humayra Tasnim, Soumya Dutta, Melanie Moses
- The research introduces a novel and adaptable method for interpreting informative features of large scale spatiotemporal data, applicable to diverse datasets from different domains. - The proposed technique identifies key informative timesteps and uses information-based fusion to summarize salient patterns of…
Keisuke Ozawa
Statistically weighted principal component analysis (wPCA) is widely used to reduce the noise of scanning transmission electron microscopy-energy-dispersive X-ray (STEM-EDX) spectroscopy data. It is beneficial to retain the spatial resolution of observation in each step of the analysis, but the direct application of…
Roberta Coletti, J. Orestes Cerdeira, Marcos Raydan, Marta B. Lopes
High-dimensional omics data often contain more variables than observations, which negatively impacts the performance of classical data analysis methods. Dimensionality reduction is typically addressed through variable selection strategies that incorporate a penalty term into the model. While effective for selecting…
Jaime Salvador–Meneses, Zoila Ruiz–Chavez, Jose Garcia–Rodriguez
The kNN (k-nearest neighbors) classification algorithm is one of the most widely used non-parametric classification methods, however it is limited due to memory consumption related to the size of the dataset, which makes them impractical to apply to large volumes of data. Variations of this method have been proposed…
Bartłomiej Fliszkiewicz, Marcin Sajdak
The aim of the following research is to assess the applicability of calculated quantum properties of molecular fragments as molecular descriptors in machine learning classification task. The research is based on bio-concentration and QM9-extended databases. A number of compounds with results from quantum-chemical…
Susanne Stoll, Elisa Infanti, Benjamin de Haas, D. Samuel Schwarzkopf
Data binning involves grouping observations into bins and calculating bin-wise summary statistics. It can cope with overplotting and noise, making it a versatile tool for comparing many observations. However, data binning goes awry if the same observations are used for binning (selection) and contrasting (selective…
Esther Heid, Charles J. McGill, Florence H. Vermeire, William H. Green
Characterizing uncertainty in machine learning models has recently gained interest in the context of machine learning reliability, robustness, safety, and active learning. Here, we separate the total uncertainty into contributions from noise in the data (aleatoric) and shortcomings of the model (epistemic), further…
Alisa Bokulich, Wendy Parker
We critically engage two traditional views of scientific data and outline a novel philosophical view that we call the pragmatic-representational (PR) view of data. On the PR view, data are representations that are the product of a process of inquiry, and they should be evaluated in terms of their adequacy or fitness…
Flore N’kam Suguem, Sébastien Déjean, Philippe Saint Pierre, Nicolas Savy
One of the challenges encountered when merging heterogeneous observational clinical datasets is the recoding of categorical target variables that may have been measured differently across data sources. Standard machine learning-based approaches, such as Multiple Imputation by Chained Equations and the k-Nearest…
Authors not listed
The recent release of Meta's Open Molecules 2025 dataset (OMol25) has enabled the creation of pretrained NNPs that can predict the energy of unseen molecules in a variety of charge and spin states. However, these models do not explicitly consider charge- or spin-based physics, potentially impacting the accuracy of…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…