27 papers · ranked by Valyu relevance
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Harsh Chhajer, Rahul Roy
Quantitative experiments are essential for investigating, uncovering and confirming our understanding of complex systems, necessitating the use of effective and robust experimental designs. Despite generally outperforming other approaches, the broader adoption of model-based design of experiments (MBDoE) has been…
Osval A. Montesinos López, Brandon Alejandro Mosqueda González, Abelardo Montesinos López, José Crossa + 2 more
'Abelardo Montesinos López' 'José Crossa' 'Nagendra Kumar Singh' 'Prasanta K. Dash'] Genomic selection (GS) is revolutionizing plant breeding. However, because it is a predictive methodology, a basic understanding of statistical machine-learning methods is necessary for its successful implementation. This methodology…
Ian Knight, Khanh Tang, John Irwin
Molecular docking is a widely used technique for leveraging protein structure in ligand discovery, but as a method, it remains difficult to utilize due to limitations that have not been adequately addressed. Despite some progress towards automation, docking still requires expert guidance, hindering its adoption by a…
Dilan Pathirana, Frank T. Bergmann, Domagoj Doresic, Polina Lakrisenko + 8 more
A central question in mathematical modeling of biological systems is determining which processes are most relevant and how they can be described. There are often competing hypotheses, which yield different models. Model comparison requires parameter optimization and sampling methods. Yet, standards for the…
Stijn Hawinkel, Olivier Thas, Steven Maere
The winner’s curse is a form of selection bias that arises when estimates are obtained for a large number of features, but only a subset of most extreme estimates is reported. It occurs in large scale significance testing as well as in rank-based selection, and imperils reproducibility of findings and follow-up study…
Authors not listed
Quantum mechanics/molecular mechanics (QM/MM) simulations are crucial for understanding enzymatic reactions, but their accuracy depends heavily on the quantum-mechanical method used. Semiempirical methods offer computational efficiency but often struggle with accuracy in complex systems. This work presents a novel…
Uroš Mlakar, Iztok Fister Jr., Iztok Fister, Heming Jia
Feature selection is essential for enhancing classification accuracy, reducing overfitting, and improving interpretability in high-dimensional datasets. Evolutionary Feature Selection (EFS) methods employ a threshold parameter $θ$ to decide feature inclusion, yet the widely used static setting $θ=0.5$ may not yield…
Authors not listed
Developing generalizable machine learning models with minimal data remains a central challenge in materials informatics. Effective models can significantly reduce costly computational simulations and time-intensive experimentation by providing reliable predictions of material properties. In this work, we investigate…
Mohamed Ghetas, Mohamed Abd Elaziz, Mohamed Issa
The presence of noisy, redundant, and irrelevant features in high-dimensional datasets significantly degrades the performance of classification models. Feature selection is a critical pre-processing step to mitigate this issue by identifying an optimal feature subset. While the Generalized Normal Distribution…
Demeke Endalie, Getamesay Haile, Wondmagegn Taye Abebe, Yilun Shang
Text classification is the process of categorizing documents based on their content into a predefined set of categories. Text classification algorithms typically represent documents as collections of words and it deals with a large number of features. The selection of appropriate features becomes important when the…
Alysson Ribeiro da Silva, Camila Guedes Silveira
—The accuracy of a classifier, when performing Pattern recognition, is mostly tied to the quality and representativeness of the input feature vector. Feature Selection is a process that allows for representing information properly and may increase the accuracy of a classifier. This process is responsible for finding…
Rahi Jain, Wei Xu
Feature selection is important in high dimensional data analysis. The wrapper approach is one of the ways to perform feature selection, but it is computationally intensive as it builds and evaluates models of multiple subsets of features. The existing wrapper approaches primarily focus on shortening the path to find an…
Zhila Yaseen Taha, Abdulhady Abas Abdullah, Tarik A. Rashid
Methods and Applications Authors: ['Zhila Yaseen Taha' 'Abdulhady Abas Abdullah' 'Tarik A. Rashid'] Analyzing large datasets to select optimal features is one of the most important research areas in machine learning and data mining. This feature selection procedure involves dimensionality reduction which is crucial in…
Gaoshuai Wang, Fabrice Lauri, Amir Hajjam El Hassani
Feature selection plays a vital role in promoting the classifier's performance. However, current methods ineffectively distinguish the complex interaction in the selected features. To further remove these hidden negative interactions, we propose a GA-like dynamic probability (GADP) method with mutual information which…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Nicolas Ngo, Pierre Michel, Roch Giorgi
\usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$$\varvec{\gamma }$$\end{document} γ -metric Authors: ['Nicolas Ngo' 'Pierre Michel' 'Roch Giorgi'] Background The…
C. Peter Sebastian, Carlos E. González-Guillén
1*Fortia Energía, Calle de Gregorio Benítez, Madrid, 28043, Spain. 2Universidad Politécnica de Madrid, Madrid, Spain. 3Departamento de Matemática Aplicada a la Ingeniería Industrial, Escuela Técnica Superior de Ingenieros Industriales, Universidad Politécnica de Madrid, Calle de José Gutiérrez Abascal, Madrid, 28006…
Firuz Kamalov, Hana Sulieman, Sherif Moussa, Jorge Avante Reyes + 1 more
'Murodbek Safaraliev'] It has been shown that while feature selection algorithms are able to distinguish between relevant and irrelevant features, they fail to differentiate between relevant and redundant and correlated features. To address this issue, we propose a highly effective approach, called Nested Ensemble…
Suruchi Jai Kumar Ahuja
A major objective of clustering is to identify groups in the data that maximizes the similarity between objects within the same cluster and minimizes the similarity between different clusters. A challenge for data clustering, and unsupervised learning in general, is that there is often no mechanism for feature…
Motahare Namakin, Modjtaba Rouhani, Mostafa Sabzekar
- As global search techniques, population-based optimization algorithms have provided promising results in feature selection (FS) problems. However, the main challenges are high time complexity due to the exploration of a large search space and consequently a large number of fitness function evaluations. Moreover, the…
Shanshan Xie, Yan Zhang, Danjv Lv, Xu Chen + 2 more
Feature selection plays a very significant role for the success of pattern recognition and data mining. Based on the maximal relevance and minimal redundancy (mRMR) method, combined with feature subset, this paper proposes an improved maximal relevance and minimal redundancy (ImRMR) feature selection method based on…
Authors not listed
This study presents a novel application of Multi-Objective Bayesian Optimization (MOBO) to enhance the formulation of flame-retardant polypropylene (PP) composites. Our goal was to optimize the chemical composition of intumescent polypropylene (PP) formulations by maximizing the Limiting Oxygen Index (LOI) and…
Fanwang Meng, Marco Martínez González, Valerii Chuiko, Alireza Tehrani + 7 more
Selector is a free, open-source Python library for selecting diverse subsets from any dataset, making it a versatile tool across a wide range of application domains. Selector implements different subset sampling algorithms based on sample distance, similarity, and spatial partitioning, along with metrics to quantify…
Ke Shang, Tianye Shu, Hisao Ishibuchi, Nan Yang + 1 more
In the evolutionary multi-objective optimization (EMO) field, the standard practice is to present the final population of an EMO algorithm as the output. However, it has been shown that the final population often includes solutions which are dominated by other solutions generated and discarded in previous generations.…
Erik D. VonKaenel, Lisa M. Bramer, Javier E. Flores, Thomas O Metz + 2 more
In recent years, high dimensional omics analyses have become more commonplace for investigating complex biological systems. Typically, these studies attempt to identify key biomolecules associated with a particular biological process. Often, machine learning (ML) is used to identify these biomolecules, typically by…
Authors not listed
The identification of kinetically feasible reaction pathways that connect a reactant to its product, including numerous intermediates and transition states, is crucial for predicting chemical reactions and elucidating reaction mechanisms. However, as molecular systems become increasingly complex or larger, the number…