27 papers · ranked by Valyu relevance
Pengyi Yang, Hao Huang, Chunlei Liu
Recent advances in single-cell biotechnologies have resulted in high-dimensional datasets with increased complexity, making feature selection an essential technique for single-cell data analysis. Here, we revisit feature selection techniques and summarise recent developments. We review their application to a range of…
Nahúm Cueto López, María Teresa García-Ordás, Facundo Vitelli-Storelli, Pablo Fernández-Navarro + 3 more
'Facundo Vitelli-Storelli' 'Pablo Fernández-Navarro' 'Camilo Palazuelos' 'Rocío Alaiz-Rodríguez' 'U Rajendra Acharya'] This study evaluates several feature ranking techniques together with some classifiers based on machine learning to identify relevant factors regarding the probability of contracting breast cancer and…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Alysson Ribeiro da Silva, Camila Guedes Silveira
—The accuracy of a classifier, when performing Pattern recognition, is mostly tied to the quality and representativeness of the input feature vector. Feature Selection is a process that allows for representing information properly and may increase the accuracy of a classifier. This process is responsible for finding…
Tomojit Ghosh, M. Kirby
Autoencoders have been widely used as a nonlinear tool for data dimensionality reduction. While autoencoders don't utilize the label information, Centroid-Encoders (CE)[1] use the class label in their learning process. In this study, we propose a sparse optimization using the Centroid-Encoder architecture to determine…
José Barrera-García, Felipe Cisternas-Caneo, Broderick Crawford, Mariam Gómez Sánchez + 2 more
Feature selection is becoming a relevant problem within the field of machine learning. The feature selection problem focuses on the selection of the small, necessary, and sufficient subset of features that represent the general set of features, eliminating redundant and irrelevant information. Given the importance of…
Muhammad Rajabinasab, Arthur Zimek
Feature selection is a fundamental machine learning and data mining task, involved with discriminating redundant features from informative ones. It is an attempt to address the curse of dimensionality by removing the redundant features, while unlike dimensionality reduction methods, preserving explainability. Feature…
Tümay Capraz, Wolfgang Huber
A fundamental step in many analyses of high-dimensional data is dimension reduction. Two basic approaches are introduction of new, synthetic coordinates, and selection of extant features. Advantages of the latter include interpretability, simplicity, transferability and modularity. A common criterion for unsupervised…
Rahi Jain, Wei Xu
Feature selection is important in high dimensional data analysis. The wrapper approach is one of the ways to perform feature selection, but it is computationally intensive as it builds and evaluates models of multiple subsets of features. The existing wrapper approaches primarily focus on shortening the path to find an…
Shanshan Xie, Yan Zhang, Danjv Lv, Xu Chen + 2 more
Feature selection plays a very significant role for the success of pattern recognition and data mining. Based on the maximal relevance and minimal redundancy (mRMR) method, combined with feature subset, this paper proposes an improved maximal relevance and minimal redundancy (ImRMR) feature selection method based on…
Atikul Islam, Tapas Bhadra, Kalyani Mali, Jayant Giri + 3 more
This article basically explores the impact of univariate & multivariate filter-based feature selection methodologies on enhancing the classification performance for the real-life classification problems. Our study considers two univariate filter-based feature selection techniques, namely, Chi-square and Fisher score…
Rahi Jain, Wei Xu
Feature selection (FS) is critical for high dimensional data analysis. Ensemble based feature selection (EFS) is a commonly used approach to develop FS techniques. Rank aggregation (RA) is an essential step of EFS where results from multiple models are pooled to estimate feature importance. However, the literature…
Jeonghwan Park, Kang Li, Huiyu Zhou
We present a new wrapper feature selection algorithm for human detection. This algorithm is a hybrid feature selection approach combining the benefits of filter and wrapper methods. It allows the selection of an optimal feature vector that well represents the shapes of the subjects in the images. In detail, the…
Annette Spooner, Gelareh Mohammadi, Perminder S. Sachdev, Henry Brodaty + 1 more
'Henry Brodaty' 'Arcot Sowmya'] Healthcare datasets present many challenges to both machine learning and statistics as their data are typically heterogeneous, censored, high-dimensional and have missing information. Feature selection is often used to identify the important features but can produce unstable results when…
Stephen Akatore Atimbire, Justice Kwame Appati, Ebenezer Owusu
Heart Diseases have the highest mortality worldwide, necessitating precise predictive models for early risk assessment. Much existing research has focused on improving model accuracy with single datasets, often neglecting the need for comprehensive evaluation metrics and utilization of different datasets in the same…
Rahi Jain, Wei Xu
Feature selection (FS) reduces the dimensions of high dimensional data. Among many FS approaches, ensemble-based feature selection (EFS) is one of the commonly used approaches. The rank aggregation (RA) step influences the feature selection of EFS. Currently, the EFS approach relies on using a single RA algorithm to…
Damir Zhakparov, Kathleen Moriarty, Damian Roqueiro, Katja Baerenfaller
High-dimensional Bulk RNA sequencing (RNAseq) datasets pose a considerable challenge in identifying biologically relevant features for downstream analyses and data mining efforts. The standard approach involves differential gene expression (DGE) analysis, but its effectiveness can be limited depending on the data due…
Nicolas Ngo, Pierre Michel, Roch Giorgi
\usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varvec{\gamma }$$\end{document} γ -metric Authors: ['Nicolas Ngo' 'Pierre Michel' 'Roch Giorgi'] Background The…
Annette Spooner, Gelareh Mohammadi, Perminder S. Sachdev, Henry Brodaty + 1 more
'Henry Brodaty' 'Arcot Sowmya' ''] Background Feature selection is often used to identify the important features in a dataset but can produce unstable results when applied to high-dimensional data. The stability of feature selection can be improved with the use of feature selection ensembles, which aggregate the…
Yongbin Zhu, Tao Li, Wenshan Li
Feature selection provides the optimal subset of features for data mining models. However, current feature selection methods for high-dimensional data also require a better balance between feature subset quality and computational cost. In this paper, an efficient hybrid feature selection method (HFIA) based on…
Muhammad Rajabinasab, Anton Danholt Lautrup, Tobias Hyrup, Arthur Zimek
'Arthur Zimek'] Abstract. Expressive evaluation metrics are indispensable for informative experiments in all areas, and while several metrics are established in some areas, in others, such as feature selection, only indirect or otherwise limited evaluation metrics are found. In this paper, we propose a novel evaluation…
Chunxu Cao, Qiang Zhang
To address this problem, we propose a novel filter feature selection method, ContrastFS, which selects discriminative features based on the discrepancies features shown between different classes. We introduce a dimensionless quantity as a surrogate representation to summarize the distributional individuality of certain…
Erik D. VonKaenel, Lisa M. Bramer, Javier E. Flores, Thomas O Metz + 2 more
In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Pulan Yu
Associative classification mining (ACM) integrating association rule mining and classification has become a significant tool for knowledge discovery, especially in the chemical domain. Its major advantage is providing high accuracy as well as chemically interpretable models. Additionally, it is able to find…
Authors not listed
Deciphering the correct mechanism governing certain phenomenon in polyelectrolyte (PE) brush grafted systems, revealed through atomistic simulations, is an extremely challenging problem. In a recent study, our all-atom molecular dynamics (MD) simulations revealed a non-linearly large electroosmotic flow (in the…
Authors not listed
This paper addresses the challenges in cell line development (CLD), the lengthy and ambiguous clone screening in upstream biopharmaceutical production. Typically, only a small subset of the later stages of CLD data is used for manually selecting lead clones. Addressing this issue, we introduce a multivariate data…