25 papers · ranked by Valyu relevance
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Nadav Harel, Tirza Routtenberg
—Statistical inference of multiple parameters often involves a preliminary parameter selection stage. The selection stage has an impact on subsequent estimation, for example by introducing a selection bias. The post-selection maximum likelihood (PSML) estimator is shown to reduce the selection bias and the…
Daniel Svensson, Rickard Sjögren, David Sundell, Andreas Sjödin + 1 more
Selecting the proper parameter settings for bioinformatic software tools is challenging. Not only will each parameter have an individual effect on the outcome, but there are also potential interaction effects between parameters. Both of these effects may be difficult to predict. To make the situation even more complex…
Ian Knight, Khanh Tang, John Irwin
Molecular docking is a widely used technique for leveraging protein structure in ligand discovery, but as a method, it remains difficult to utilize due to limitations that have not been adequately addressed. Despite some progress towards automation, docking still requires expert guidance, hindering its adoption by a…
Ian Knight, Khanh Tang, Olivier Mailhot, John Irwin
Molecular docking is a widely used technique for leveraging protein structure in ligand discovery, but as a method, it remains difficult to utilize due to limitations that have not been adequately addressed. Despite some progress towards automation, docking still requires expert guidance, hindering its adoption by a…
Dilan Pathirana, Frank T. Bergmann, Domagoj Doresic, Polina Lakrisenko + 8 more
A central question in mathematical modeling of biological systems is determining which processes are most relevant and how they can be described. There are often competing hypotheses, which yield different models. Model comparison requires parameter optimization and sampling methods. Yet, standards for the…
Authors not listed
Quantum mechanics/molecular mechanics (QM/MM) simulations are crucial for understanding enzymatic reactions, but their accuracy depends heavily on the quantum-mechanical method used. Semiempirical methods offer computational efficiency but often struggle with accuracy in complex systems. This work presents a novel…
Rina Onda, Zheng‐Yan Gao, Masaaki Kotera, Kenta Oono
It is preferred that feature selectors be stable for better interpretabity and robust prediction. Ensembling is known to be effective for improving the stability of feature selectors. Since ensembling is time-consuming, it is desirable to reduce the computational cost to estimate the stability of the ensemble feature…
Uroš Mlakar, Iztok Fister Jr., Iztok Fister, Heming Jia
Feature selection is essential for enhancing classification accuracy, reducing overfitting, and improving interpretability in high-dimensional datasets. Evolutionary Feature Selection (EFS) methods employ a threshold parameter $θ$ to decide feature inclusion, yet the widely used static setting $θ=0.5$ may not yield…
Karan Uppal, Eva K. Lee
Recent studies have shown that the ensemble feature selection approaches are essential for generating robust classifiers. Existing methods for aggregating feature lists from different methods require use of arbitrary thresholds for selecting the top ranked features and do not account for classification accuracy while…
Yosef Masoudi-Sobhanzadeh, Habib Motieghader, Ali Masoudi-Nejad
Background Feature selection, as a preprocessing stage, is a challenging problem in various sciences such as biology, engineering, computer science, and other fields. For this purpose, some studies have introduced tools and softwares such as WEKA. Meanwhile, these tools or softwares are based on filter methods which…
Zhila Yaseen Taha, Abdulhady Abas Abdullah, Tarik A. Rashid
Methods and Applications Authors: ['Zhila Yaseen Taha' 'Abdulhady Abas Abdullah' 'Tarik A. Rashid'] Analyzing large datasets to select optimal features is one of the most important research areas in machine learning and data mining. This feature selection procedure involves dimensionality reduction which is crucial in…
C. Fernandez-Lozano, C. Canto, M. Gestal, J. M. Andrade-Garda + 3 more
'J. R. Rabuñal' 'J. Dorado' 'A. Pazos'] Given the background of the use of Neural Networks in problems of apple juice classification, this paper aim at implementing a newly developed method in the field of machine learning: the Support Vector Machines (SVM). Therefore, a hybrid model that combines genetic algorithms…
Authors not listed
Developing generalizable machine learning models with minimal data remains a central challenge in materials informatics. Effective models can significantly reduce costly computational simulations and time-intensive experimentation by providing reliable predictions of material properties. In this work, we investigate…
Mohamed Ghetas, Mohamed Abd Elaziz, Mohamed Issa
The presence of noisy, redundant, and irrelevant features in high-dimensional datasets significantly degrades the performance of classification models. Feature selection is a critical pre-processing step to mitigate this issue by identifying an optimal feature subset. While the Generalized Normal Distribution…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Nicolas Ngo, Pierre Michel, Roch Giorgi
\usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varvec{\gamma }$$\end{document} γ -metric Authors: ['Nicolas Ngo' 'Pierre Michel' 'Roch Giorgi'] Background The…
Demeke Endalie, Getamesay Haile, Wondmagegn Taye Abebe, Yilun Shang
Text classification is the process of categorizing documents based on their content into a predefined set of categories. Text classification algorithms typically represent documents as collections of words and it deals with a large number of features. The selection of appropriate features becomes important when the…
Motahare Namakin, Modjtaba Rouhani, Mostafa Sabzekar
- As global search techniques, population-based optimization algorithms have provided promising results in feature selection (FS) problems. However, the main challenges are high time complexity due to the exploration of a large search space and consequently a large number of fitness function evaluations. Moreover, the…
Deniz Akdemir
Optimal subset selection is an important task that has numerous algorithms designed for it and has many application areas. STPGA contains a special genetic algorithm supplemented with a tabu memory property (that keeps track of previously tried solutions and their fitness for a number of iterations), and with a…
Authors not listed
This study presents a novel application of Multi-Objective Bayesian Optimization (MOBO) to enhance the formulation of flame-retardant polypropylene (PP) composites. Our goal was to optimize the chemical composition of intumescent polypropylene (PP) formulations by maximizing the Limiting Oxygen Index (LOI) and…
Faizal Hafiz, Akshya Swain, Nitish Patel, Chirag A. Naik
This paper proposes a new generalized two dimensional learning approach for particle swarm based feature selection. The core idea of the proposed approach is to include the information about the subset cardinality into the learning framework by extending the dimension of the velocity. The 2D-learning framework retains…
Alysson Ribeiro da Silva, Camila Guedes Silveira
—The accuracy of a classifier, when performing Pattern recognition, is mostly tied to the quality and representativeness of the input feature vector. Feature Selection is a process that allows for representing information properly and may increase the accuracy of a classifier. This process is responsible for finding…
Filip Koprivec, Klemen Kenda, Beno Šircelj
In this paper, a novel feature selection algorithm for inference from high-dimensional data (FASTENER) is presented. With its multi-objective approach, the algorithm tries to maximize the accuracy of a machine learning algorithm with as few features as possible. The algorithm exploits entropy-based measures, such as…
Ke Shang, Tianye Shu, Hisao Ishibuchi, Nan Yang + 1 more
In the evolutionary multi-objective optimization (EMO) field, the standard practice is to present the final population of an EMO algorithm as the output. However, it has been shown that the final population often includes solutions which are dominated by other solutions generated and discarded in previous generations.…