23 papers · ranked by Valyu relevance
Ursula Neumann, Nikita Genze, Dominik Heider
Background Feature selection methods aim at identifying a subset of features that improve the prediction performance of subsequent classification models and thereby also simplify their interpretability. Preceding studies demonstrated that single feature selection methods can have specific biases, whereas an ensemble…
Annette Spooner, Gelareh Mohammadi, Perminder S. Sachdev, Henry Brodaty + 1 more
'Henry Brodaty' 'Arcot Sowmya' ''] Background Feature selection is often used to identify the important features in a dataset but can produce unstable results when applied to high-dimensional data. The stability of feature selection can be improved with the use of feature selection ensembles, which aggregate the…
Ursula Neumann, Mona Riemenschneider, Jan-Peter Sowa, Theodor Baars + 3 more
'Julia Kälsch' 'Ali Canbay' 'Dominik Heider'] Motivation Biomarker discovery methods are essential to identify a minimal subset of features (e.g., serum markers in predictive medicine) that are relevant to develop prediction models with high accuracy. By now, there exist diverse feature selection methods, which either…
Tao He, Jason Min Baik, Chiemi Kato, Hai Yang + 3 more
The T and B cell repertoire make up the adaptive immune system and is mainly generated through somatic V(D)J gene recombination. Thus, the VJ gene usage may be a potential prognostic or predictive biomarker. However, analysis of the adaptive immune system is challenging due to the heterogeneity of the clonotypes that…
Pijush Das, Anirban Roychowdhury, Subhadeep Das, Susanta Roychoudhury + 1 more
'Susanta Roychoudhury' 'Sucheta Tripathy'] Biological data are accumulating at a faster rate, but interpreting them still remains a problem. Classifying biological data into distinct groups is the first step in understanding them. Data classification in response to a certain treatment is an extremely important aspect…
Annette Spooner, Gelareh Mohammadi, Perminder S. Sachdev, Henry Brodaty + 1 more
'Henry Brodaty' 'Arcot Sowmya'] Healthcare datasets present many challenges to both machine learning and statistics as their data are typically heterogeneous, censored, high-dimensional and have missing information. Feature selection is often used to identify the important features but can produce unstable results when…
Xiaokang Zhang, Inge Jonassen
Ensemble learning that can be used to combine the predictions from multiple learners has been widely applied in pattern recognition, and has been reported to be more robust and accurate than the individual learners. This ensemble logic has recently also been more applied in feature selection. There are basically two…
Rahi Jain, Wei Xu
Feature selection (FS) reduces the dimensions of high dimensional data. Among many FS approaches, ensemble-based feature selection (EFS) is one of the commonly used approaches. The rank aggregation (RA) step influences the feature selection of EFS. Currently, the EFS approach relies on using a single RA algorithm to…
Karan Uppal, Eva K. Lee
Recent studies have shown that the ensemble feature selection approaches are essential for generating robust classifiers. Existing methods for aggregating feature lists from different methods require use of arbitrary thresholds for selecting the top ranked features and do not account for classification accuracy while…
Rahi Jain, Wei Xu
Feature selection (FS) is critical for high dimensional data analysis. Ensemble based feature selection (EFS) is a commonly used approach to develop FS techniques. Rank aggregation (RA) is an essential step of EFS where results from multiple models are pooled to estimate feature importance. However, the literature…
Ali Anaissi, Madhu Goyal, Daniel R. Catchpoole, Ali Braytee + 2 more
'Paul J. Kennedy' 'Bin Liu'] The identification of a subset of genes having the ability to capture the necessary information to distinguish classes of patients is crucial in bioinformatics applications. Ensemble and bagging methods have been shown to work effectively in the process of gene selection and classification.…
Zekun Xin, Ruhong Lv, Wei Liu, Shenghan Wang + 4 more
'Guangyu Sun' 'Xiangtao Li'] Feature selection plays a crucial role in classification tasks as part of the data preprocessing process. Effective feature selection can improve the robustness and interpretability of learning algorithms, and accelerate model learning. However, traditional statistical methods for feature…
Li-Hsin Cheng, Che Lin
Breast cancer is a heterogeneous disease. In order to guide proper treatment decisions for each individual patient, there is an urgent need for robust prognostic biomarkers that allow reliable prognosis prediction. Gene feature selection on microarray data is an approach to systematically discover potential biomarkers.…
Annette Spooner, Gelareh Mohammadi, Perminder S. Sachdev, Henry Brodaty + 1 more
'Henry Brodaty' 'Arcot Sowmya'] Healthcare datasets often contain groups of highly correlated features, such as features from the same biological system. When feature selection is applied to these datasets to identify the most important features, the biases inherent in some multivariate feature selectors due to…
Yunpu Zhao
Automated machine learning has achieved remarkable technological developments in recent years, and building an automated machine learning pipeline is now an essential task. However, existing AutoML pipeline approaches adopt monotonous ensemble strategies across different machine learning classification tasks. They…
Michela Carlotta Massi, Francesca Gasperoni, Francesca Ieva, Anna Maria Paganoni
'Anna Maria Paganoni'] Class imbalance is a common issue in many domain applications of learning algorithms. Oftentimes, in the same domains it is much more relevant to correctly classify and profile minority class observations. This need can be addressed by Feature Selection (FS), that offers several further…
H.M.Fazlul Haque, Fariha Arifin, Sheikh Adilina, Muhammod Rafsanjani + 1 more
The information of a cell is primarily contained in Deoxyribonucleic Acid (DNA). There is a flow of information of DNA to protein sequences via Ribonucleic acids (RNA) through transcription and translation. These entities are vital for the genetic process. Recent developments in epigenetic also show the importance of…
Sebastian Spänig, Alexander Michel, Dominik Heider
Owing to the rising levels of multi-resistant pathogens, antimicrobial peptides, an alternative strategy to classic antibiotics, got more attention. A crucial part is thereby the costly identification and validation. With the ever-growing amount of annotated peptides, researchers employed artificial intelligence to…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Authors not listed
We present a multi-stage framework for predictive modeling that integrates automated feature engineering, selective dimensionality reduction, and targeted ensembling. Our pipeline begins with feature generation using a GPU-accelerated adaptation of AutoFeat, followed by variance-based pruning and LightGBM gain-based…
Authors not listed
Machine learning holds significant promise for accelerating biomarker discovery in clinical proteomics, yet its real-world impact remains limited by widespread methodological pitfalls and unrealistic expectations. In this perspective, we critically examine the integration of machine learning into clinical proteomics…
Moayad Alnammi, Shengchao Liu, Spencer S Ericksen, Gene E Ananiev + 6 more
Traditional small molecule drug discovery is a time consuming and costly endeavor. High-throughput chemical screening can only assess a tiny fraction of drug-like chemical space. The strong predictive power of modern machine learning methods for virtual chemical screening enables training models on known active and…
Authors not listed
Ensuring the trustworthiness of machine learning (ML) models in high-stake applications is crucial. One such application is predicting anti-cancer drug sensitivity, where ML models are built with the final goal of integrating them into treatment recommendation systems for personalized medicine. Here, we propose a…