30 papers · ranked by Valyu relevance
Carl H Lubba, Sarab S Sethi, Philip Knaute, Simon R Schultz + 2 more
Capturing the dynamical properties of time series concisely as interpretable feature vectors can enable efficient clustering and classification for time-series applications across science and industry. Selecting an appropriate feature-based representation of time series for a given application can be achieved through…
Jundong Li, Kewei Cheng, Suhang Wang, Fred Morstatter + 3 more
'Robert P. Treviño' 'Jiliang Tang' 'Huan Liu'] Feature selection, as a data preprocessing strategy, has been proven to be effective and efficient in preparing data (especially high-dimensional data) for various data mining and machine learning problems. The objectives of feature selection include: building simpler and…
Yosef Masoudi-Sobhanzadeh, Habib Motieghader, Ali Masoudi-Nejad
Background Feature selection, as a preprocessing stage, is a challenging problem in various sciences such as biology, engineering, computer science, and other fields. For this purpose, some studies have introduced tools and softwares such as WEKA. Meanwhile, these tools or softwares are based on filter methods which…
Authors not listed
Traditional and non-classical machine learning models for solid-state structure prediction have predominantly relied on compositional features (derived from properties of constituent elements) to predict the existence of structure and its properties. However, the lack of structural information can be a source of…
Tümay Capraz, Wolfgang Huber
A fundamental step in many analyses of high-dimensional data is dimension reduction. Two basic approaches are introduction of new, synthetic coordinates, and selection of extant features. Advantages of the latter include interpretability, simplicity, transferability and modularity. A common criterion for unsupervised…
Michelle R. Greene, Bruce C. Hansen
Human scene categorization is characterized by its remarkable speed. While many visual and conceptual features have been linked to this ability, significant correlations exist between feature spaces, impeding our ability to determine their relative contributions to scene categorization. Here, we employed a whitening…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
René-Vinicio Sánchez, Jean Carlo Macancela, Luis-Renato Ortega, Diego Cabrera + 3 more
'Diego Cabrera' 'Fausto Pedro García Márquez' 'Mariela Cerrada' 'Jiawei Xiang'] This article presents a comprehensive collection of formulas and calculations for hand-crafted feature extraction of condition monitoring signals. The documented features include 123 for the time domain and 46 for the frequency domain.…
Suguru Fujita, Yasuaki Karasawa, Ken-ichi Hironaka, Y-h. Taguchi + 1 more
High-throughput omics technologies have enabled the profiling of entire biological systems. For the biological interpretation of such omics data, two analyses, hypothesis- and data-driven analyses including tensor decomposition, have been used. Both analyses have their own advantages and disadvantages and are mutually…
Tianping Zhang, Zheyu Zhang, Zhiyuan Fan, Haoyan Luo + 3 more
'Wei Cao' 'Jian Li'] The goal of automated feature generation is to liberate machine learning experts from the laborious task of manual feature generation, which is crucial for improving the learning performance of tabular data. The major challenge in automated feature generation is to efficiently and accurately…
Terry N. Guo, Animesh Dahal, Ambareen Siraj
This paper presents analytical techniques to improve redundancy and relevance assessment for precise selection of features in practical multi-class raw datasets. We propose a matrix-rank based k-medoids algorithm that guarantees to output all independent medoids. The new algorithm uses matrix rank as a robust…
Rahi Jain, Wei Xu
Feature selection (FS) reduces the dimensions of high dimensional data. Among many FS approaches, ensemble-based feature selection (EFS) is one of the commonly used approaches. The rank aggregation (RA) step influences the feature selection of EFS. Currently, the EFS approach relies on using a single RA algorithm to…
Rahi Jain, Wei Xu
Feature selection (FS) is critical for high dimensional data analysis. Ensemble based feature selection (EFS) is a commonly used approach to develop FS techniques. Rank aggregation (RA) is an essential step of EFS where results from multiple models are pooled to estimate feature importance. However, the literature…
Xiongshi Deng, Min Li, Lei Wang, Qikang Wan
—Feature selection is a preprocessing step which plays a crucial role in the domain of machine learning and data mining. Feature selection methods have been shown to be effective in removing redundant and irrelevant features, improving the learning algorithm's prediction performance. Among the various methods of…
Sangjoon Lee, Clio Chen, Griheydi Garcia, Anton Oliynyk
Materials informatics uses data-driven approaches for the study and discovery of materials. Features or descriptors are the crucial components in generating reliable and accurate machine-learning models. While general data can be acquired through public and commercial sources, features must be tailored for a specific…
Khalil Taheri, Hadi Moradi, Mostafa Tavassolipour
One of the most important problems in the field of pattern recognition is data classification. Due to the increasing development of technologies introduced in the field of data classification, some of the solutions are still open and need more research. One of the challenging problems in this area is the curse of…
Snehasis Banerjee, Tanushyam Chattopadhyay, A. Mukherjee
This paper presents an automated approach for interpretable feature recommendation for solving signal data analytics problems. The method has been tested by performing experiments on datasets in the domain of prognostics where interpretation of features is considered very important. The proposed approach is based on…
Xiaohong Wang, Yidi He, Lizhi Wang
In this study, due to the redundant and irrelevant features contained in the multi-dimensional feature parameter set, the information fusion performance of the subspace learning algorithm was reduced. To solve the above problem, a mutual information (MI) and fractal dimension-based unsupervised feature parameters…
Suvo Banik, Karthik Balasubramanian, Sukriti Manna, Sybil Derrible + 1 more
Identifying key descriptors and understanding important features across different classes of materials are crucial for machine learning (ML) tools to both predict material properties and reveal the physics underlying any process of interest. Traditionally, the predictive modeling of elastic properties of materials is…
Bhavana R. Bhamare, Jeyanthi Prabhu, Sebastian Ventura
Due to the massive progression of the Web, people post their reviews for any product, movies and places they visit on social media. The reviews available on social media are helpful to customers as well as the product owners to evaluate their products based on different reviews. Analyzing structured data is easy as…
Reem Salman, Ayman Alzaatreh, Hana Sulieman, Shaimaa Faisal + 1 more
'Mohamed Medhat Gaber'] In the past decade, big data has become increasingly prevalent in a large number of applications. As a result, datasets suffering from noise and redundancy issues have necessitated the use of feature selection across multiple domains. However, a common concern in feature selection is that…
Steven Torrisi, Matthew Carbone, Brian Rohr, Joseph H. Montoya + 4 more
X-ray absorption spectroscopy (XAS) produces a wealth of information about the local structure of materials, but interpretation of spectra often relies on easily accessible trends and prior assumptions about the structure. Recently, researchers have demonstrated that machine learning models can automate this process to…
Urszula Stańczyk
The performance of a classification system of any type can suffer from irrelevant or redundant data, contained in characteristic features that describe objects of the universe. To estimate relevance of attributes and select their subset for a constructed classifier typically either a filter, wrapper, or an embedded…
Lior Friedman, Shaul Markovitch
When humans perform inductive learning, they often enhance the process with background knowledge. With the increasing availability of well-formed collaborative knowledge bases, the performance of learning algorithms could be significantly enhanced if a way were found to exploit these knowledge bases. In this work, we…
Muhammad Rajabinasab, Anton Danholt Lautrup, Tobias Hyrup, Arthur Zimek
'Arthur Zimek'] Abstract. Expressive evaluation metrics are indispensable for informative experiments in all areas, and while several metrics are established in some areas, in others, such as feature selection, only indirect or otherwise limited evaluation metrics are found. In this paper, we propose a novel evaluation…
Fatima Skaka-Čekić, Jasmina Baraković Husić, Almasa Odžak, Mesud Hadžialić + 2 more
Big Data analytics and Artificial Intelligence (AI) technologies have become the focus of recent research due to the large amount of data. Dimensionality reduction techniques are recognized as an important step in these analyses. The multidimensional nature of Quality of Experience (QoE) is based on a set of Influence…
Soukhin Das, G.R. Mangun, Mingzhou Ding
Perceptual expertise and attention are two important factors that enable superior object recognition and task performance. While expertise enhances knowledge and provides a holistic understanding of the environment, attention allows us to selectively focus on task-related information and suppress distraction. It has…
Authors not listed
The global drive towards net-zero has accelerated the adoption of carbon fibre reinforced polymers (CFRP) for lightweight structures in various sectors such as aerospace, automotive, energy and biomedical. Mechanical machining of CFRP is often necessary to meet dimensional or assembly-related requirements. However…
Anthony Onwuli, Keith T. Butler, Aron Walsh
High-dimensional representations of the elements have become common within the field of materials informatics to build useful, structure-agnostic models for the chemistry of materials. However, the characteristics of elements change when they adopt a given oxidation state, with distinct structural preferences and…
Adamantios Ntakaris, Juho Kanniainen, Moncef Gabbouj, Alexandros Iosifidis + 1 more
'Alexandros Iosifidis' 'Alejandro Raul Hernandez Montoya'] Stock price prediction is a challenging task, in which machine learning methods have recently been successfully used. In this paper, we extract over 270 hand-crafted features (factors) inspired by technical indicators and quantitative analysis and test their…