Search · four archives
Search · four archives
27 papers · ranked by Valyu relevance
Jundong Li, Kewei Cheng, Suhang Wang, Fred Morstatter + 3 more
'Robert P. Treviño' 'Jiliang Tang' 'Huan Liu'] Feature selection, as a data preprocessing strategy, has been proven to be effective and efficient in preparing data (especially high-dimensional data) for various data mining and machine learning problems. The objectives of feature selection include: building simpler and…
Zhaozhao Xu, Fangyuan Yang, Hong Wang, Junding Sun + 3 more
2## Related works Unsupervised feature selection (; ; ) reduces the dimensionality of data by mapping data from a high-dimensional space to a low-dimensional space and removing redundant and irrelevant features. In this section, we provide an overview of classic unsupervised feature selection algorithms, which can be…
Wei Liu, Qian Ning, Guangwei Liu, Haonan Wang + 3 more
'Miao Zhong' 'Khan Bahadar Khan'] Traditional subspace feature selection methods typically rely on a fixed distance to compute residuals between the original and feature reconstruction spaces. However, this approach struggles to adapt to diverse datasets and often fails to handle noise and outliers effectively. In this…
Assaf Gottlieb, Roy Varshavsky, Michal Linial, David Horn
Background Feature selection is an important pre-processing task in the analysis of complex data. Selecting an appropriate subset of features can improve classification or clustering and lead to better understanding of the data. An important example is that of finding an informative group of genes out of thousands that…
Min Yan, Mao Ye, Tian Liang, Yulin Jian + 2 more
Feature selection is a widely used dimension reduction technique to select feature subsets because of its interpretability. Many methods have been proposed and achieved good results, in which the relationships between adjacent data points are mainly concerned. But the possible associations between data pairs that are…
Juanying Xie, Mingzhao Wang, Shengquan Xu, Zhao Huang + 1 more
'Philip W. Grant'] To tackle the challenges in genomic data analysis caused by their tens of thousands of dimensions while having a small number of examples and unbalanced examples between classes, the technique of unsupervised feature selection based on standard deviation and cosine similarity is proposed in this…
Ziheng Sun, Chris Ding, Jicong Fan
—Feature selection is important for high-dimensional data analysis and is non-trivial in unsupervised learning problems such as dimensionality reduction and clustering. The goal of unsupervised feature selection is finding a subset of features such that the data points from different clusters are well separated. This…
Zhenzhen Sun, Yuanlong Yu
Feature selection is an important data preprocessing in data mining and machine learning which can be used to reduce the feature dimension without deteriorating model's performance. Since obtaining annotated data is laborious or even infeasible in many cases, unsupervised feature selection is more practical in reality.…
Peican Zhu, Xin Hou, Zhen Wang, Feiping Nie
Along with the flourish of the information age, massive amounts of data are generated day by day. Due to the large-scale and high-dimensional characteristics of these data, it is often difficult to achieve better decision-making in practical applications. Therefore, an efficient big data analytics method is urgently…
Gang Chen, Yuanli Cai, Juan Shi
Feature selection, also known as attribute selection, is the technique of selecting a subset of relevant features for building robust object models. It is becoming more and more important for large-scale sensors applications with AI capabilities. The core idea of this paper is derived from a straightforward and…
Y-h. Taguchi, Turki Turki
Identifying differentially expressed genes is difficult because of the small number of available samples compared with the large number of genes. Conventional gene selection methods employing statistical tests have the critical problem of heavy dependence of P-values on sample size. Although the recently proposed…
Sen Wang, Feiping Nie, Xiaojun Chang, Lina Yao + 2 more
'Quan Z. Sheng'] Abstract. Unsupervised feature selection has been always attracting research attention in the communities of machine learning and data mining for decades. In this paper, we propose an unsupervised feature selection method seeking a feature coefficient matrix to select the most distinctive features.…
Cihan Kuzudisli, Burcu Bakir-Gungor, Nurten Bulut, Bahjat Qaqish + 2 more
'Malik Yousef' 'Reema Singh'] With the rapid development in technology, large amounts of high-dimensional data have been generated. This high dimensionality including redundancy and irrelevancy poses a great challenge in data analysis and decision making. Feature selection (FS) is an effective way to reduce…
Sriparna Saha, Asif Ekbal, Abhay Kumar Alok, Rachamadugu Spandana
In this paper we have coupled feature selection problem with semi-supervised clustering. Semi-supervised clustering utilizes the information of unsupervised and supervised learning in order to overcome the problems related to them. But in general all the features present in the data set may not be important for…
Tümay Capraz, Wolfgang Huber
A fundamental step in many analyses of high-dimensional data is dimension reduction. Two basic approaches are introduction of new, synthetic coordinates, and selection of extant features. Advantages of the latter include interpretability, simplicity, transferability and modularity. A common criterion for unsupervised…
Y-h. Taguchi, Turki Turki
In this work, we extended the recently developed tensor decomposition (TD) based unsupervised feature extraction (FE) to a kernel based method, through a mathematical formulation. Subsequently, the kernel TD (KTD) based unsupervised FE was applied to two synthetic examples as well as real data sets, and the relevant…
YH. Taguchi, Turki Turki
Feature selection of multi-omics data analysis remains challenging since omics data include 10^2^–10^5^ features. How to weight an individual omics dataset is unclear and greatly affects feature selection consequences. In this study, a recently proposed kernel tensor decomposition (KTD)-based unsupervised feature…
Suruchi Jai Kumar Ahuja
A major objective of clustering is to identify groups in the data that maximizes the similarity between objects within the same cluster and minimizes the similarity between different clusters. A challenge for data clustering, and unsupervised learning in general, is that there is often no mechanism for feature…
Ngan Thi Dong, Megha Khosla
The identification of biomarkers or predictive features that are indicative of a specific biological or disease state is a major research topic in biomedical applications. Several feature selection(FS) methods ranging from simple univariate methods to recent deep-learning methods have been proposed to select a minimal…
Authors not listed
X-ray diffraction (XRD) is an immediate and powerful characterization technique that provides detailed information on the lattice structure and long-range order in crystalline materials. In recent decades, the quality and quantity of available crystal structure data has exploded, in large part due to the advent of…
Giorgio Roffo, Simone Melzi
In an era where accumulating data is easy and storing it inexpensive, feature selection plays a central role in helping to reduce the high-dimensionality of huge amounts of otherwise meaningless data. In this paper, we propose a graph-based method for feature selection that ranks features by identifying the most…
Authors not listed
The analysis of nonadiabatic molecular dynamics (NAMD) data presents significant challenges due to its high dimensionality and complexity. To address these issues, we introduce ULaMDyn, a Python-based, open-source package designed to automate the unsupervised analysis of large datasets generated by NAMD simulations.…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Authors not listed
The discovery of chemically novel or structurally anomalous metal-organic frameworks (MOFs) is essential for expanding reticular design space and enhancing dataset reliability. We present CHEM-AD (Chemically Unusual Metal–organic Frameworks via Autoencoder-based Detection), a label-free, CPU-efficient pipeline that…
Authors not listed
Acoustic measurements of batteries are known to be correlated to their state-of-charge, creating opportunities for state estimation that do not rely on electrical signals. State estimators are typically parametric models fitted from data, often from the broad toolbox of machine learning. Such models can be easily…
Authors not listed
Deciphering the correct mechanism governing certain phenomenon in polyelectrolyte (PE) brush grafted systems, revealed through atomistic simulations, is an extremely challenging problem. In a recent study, our all-atom molecular dynamics (MD) simulations revealed a non-linearly large electroosmotic flow (in the…
Pulan Yu
Associative classification mining (ACM) integrating association rule mining and classification has become a significant tool for knowledge discovery, especially in the chemical domain. Its major advantage is providing high accuracy as well as chemically interpretable models. Additionally, it is able to find…