25 papers · ranked by Valyu relevance
Aparna Nair-Kanneganti, Trevor Chan, Shir Goldfinger, Emily Mackay + 2 more
Despite huge advances, LLMs still lack convenient and reliable methods to quantify the uncertainty in their responses, making them difficult to trust in high-stakes applications. One of the simplest approaches to eliciting more accurate answers is to select the mode of many responses, a technique known as ensembling.…
Rabea Khatun, Maksuda Akter, Md. Manowarul Islam, Md. Ashraf Uddin + 8 more
Biomarker-based cancer identification and classification tools are widely used in bioinformatics and machine learning fields. However, the high dimensionality of microarray gene expression data poses a challenge for identifying important genes in cancer diagnosis. Many feature selection algorithms optimize cancer…
Nikolaos Peppes, Emmanouil Daskalakis, Theodoros Alexakis, Evgenia Adamopoulou + 2 more
'Evgenia Adamopoulou' 'Konstantinos Demestichas' 'Ivan Andonovic'] The upcoming agricultural revolution, known as Agriculture 4.0, integrates cutting-edge Information and Communication Technologies in existing operations. Various cyber threats related to the aforementioned integration have attracted increasing interest…
Cristina Cornelio, Michele Donini, Andrea Loreggia, Maria Pini + 1 more
'Francesca Rossi'] Abstract In many machine learning scenarios, looking for the best classifier that fits a particular dataset can be very costly in terms of time and resources. Moreover, it can require deep knowledge of the specific domain. We propose a new technique which does not require profound expertise in the…
Eric Bax
For a voting ensemble that selects an odd-sized subset of the ensemble classifiers at random for each example, applies them to the example, and returns the majority vote, we show that any number of voters may minimize the error rate over an out-of-sample distribution. The optimal number of voters depends on the…
Yong-Woon Kim, Yung-Cheol Byun, Addapalli V. N. Krishna, Chun-Hung Liu
'Chun-Hung Liu'] Image segmentation plays a central role in a broad range of applications, such as medical image analysis, autonomous vehicles, video surveillance and augmented reality. Portrait segmentation, which is a subset of semantic image segmentation, is widely used as a preprocessing step in multiple…
Sofia Escudero, Sofia Duarte, Rosario Vitale, Emilio Fenoy + 3 more
Due to the rapid growth of sequence generation, which has surpassed the expert curators ability to manually review and annotate them, the computational annotation of proteins remains a significant challenge in bioinformatics nowadays. The Pfam database contains a large collection of proteins that are nowadays annotated…
Hayden Chen
The identification of an effective inhibitor is an essential starting point in drug discovery. Unfortunately, many issues arise with conventional high-throughput screening methods. Thus, new strategies are needed to filter through large compound screening libraries to create target-focused, smaller libraries. Effective…
DongSeong-Yoon
—Since the Fourth Industrial Revolution, AI technology has been widely used in many fields, but there are several limitations that need to be overcome, including overfitting/underfitting, class imbalance, and the limitations of representation (hypothesis space) due to the characteristics of different models. As a…
Maryam Sabzevari, Gonzalo Martínez-Muñoz, Alberto Suárez
Vote-boosting is a sequential ensemble learning method in which the individual classifiers are built on different weighted versions of the training data. To build a new classifier, the weight of each training instance is determined in terms of the degree of disagreement among the current ensemble predictions for that…
Robert D Clark, Wenkel Liang, Adam C Lee, Michael S Lawless + 2 more
'Robert Fraczkiewicz' 'Marvin Waldman'] Background Quantitative structure-activity (QSAR) models have enormous potential for reducing drug discovery and development costs as well as the need for animal testing. Great strides have been made in estimating their overall reliability, but to fully realize that potential…
Julien Knafou, Quentin Haas, Nikolay Borissov, Michel Counotte + 7 more
The COVID-19 pandemic has led to an unprecedented amount of scientific publications, growing at a pace never seen before. Multiple living systematic reviews have been developed to assist professionals with up-to-date and trustworthy health information, but it is increasingly challenging for systematic reviewers to keep…
Shengli Wu, Weimin Ding
Ensemble classifiers have been investigated by many in the artificial intelligence and machine learning community. Majority voting and weighted majority voting are two commonly used combination schemes in ensemble learning. However, understanding of them is incomplete at best, with some properties even misunderstood.…
Palak Mahajan, Shahadat Uddin, Farshid Hajati, Mohammad Ali Moni + 1 more
'Joaquim Carreras'] Machine learning models are used to create and enhance various disease prediction frameworks. Ensemble learning is a machine learning technique that combines multiple classifiers to improve performance by making more accurate predictions than a single classifier. Although numerous studies have…
Ling Luo, Po-Ting Lai, Chih-Hsuan Wei, Zhiyong Lu
Automatic extracting interactions between chemical compound/drug and gene/protein is significantly beneficial to drug discovery, drug repurposing, drug design, and biomedical knowledge graph construction. To promote the development and evaluation of systems that are able to automatically detect in relations between…
Xuan-Truc Dinh Tran, Tieu-Long Phan, Van-Thinh To, Ngoc-Vi Nguyen Tran + 4 more
3D pharmacophore models describe the ligand’s chemical interactions in their bioactive conformation. They offer a simple but sophisticated approach to decipher the chemically encoded ligand information, making them a valuable tool in Drug Design. Our research summarized the key studies for applying 3D pharmacophore…
Maryam Mahsal Khan, Alexandre Mendes, Stephan K. Chalup, Pratyoosh Shukla
'Pratyoosh Shukla'] Wavelet Neural Networks are a combination of neural networks and wavelets and have been mostly used in the area of time-series prediction and control. Recently, Evolutionary Wavelet Neural Networks have been employed to develop cancer prediction models. The present study proposes to use ensembles of…
Erdal Tasci, Ying Zhuge, Harpreet Kaur, Kevin Camphausen + 2 more
'Andra Valentina Krauze' 'Lorenzo Corsi'] Determining the aggressiveness of gliomas, termed grading, is a critical step toward treatment optimization to increase the survival rate and decrease treatment toxicity for patients. Streamlined grading using molecular information has the potential to facilitate decision…
Shehzad Khalid, Sannia Arshad, Sohail Jabbar, Seungmin Rho
We have presented a classification framework that combines multiple heterogeneous classifiers in the presence of class label noise. An extension of m-Mediods based modeling is presented that generates model of various classes whilst identifying and filtering noisy training data. This noise free data is further used to…
Esther Heid, Charles J. McGill, Florence H. Vermeire, William H. Green
Characterizing uncertainty in machine learning models has recently gained interest in the context of machine learning reliability, robustness, safety, and active learning. Here, we separate the total uncertainty into contributions from noise in the data (aleatoric) and shortcomings of the model (epistemic), further…
Authors not listed
In molecular machine learning, the choice of the representation of molecules can have a significant impact on model performance. However, understanding the root causes of these performance differences often proves challenging. One promising approach to explore model behavior is representational alignment, which…
Authors not listed
This study presents a validation and refinement of the “yellow cards” error detection workflow that can be applied to any property connected to molecular structure. In our implementation the workflow employed 5 predictive models with each assigning a “yellow card” to 5% of the entries with worst prediction accuracy.…
Agastya P Bhati, Peter V. Coveney
The accurate and reliable prediction of protein-ligand binding affinities can play a central role in the drug discovery process as well as in personalised medicine. Of considerable importance during lead optimisation are the alchemical free energy methods that furnish estimation of relative binding free energies (RBFE)…
Jürgen Köfinger, Gerhard Hummer
The proper balancing of information from experiment and theory is a long-standing problem in the analysis of noisy and incomplete data. Viewed as a Pareto optimization problem, improved agreement with the experimental data comes at the expense of growing inconsistencies with the theoretical reference model. Here, we…
Jürgen Köfinger, Gerhard Hummer
The proper balancing of information from experiment and theory is a long-standing problem in the analysis of noisy and incomplete data. Viewed as a Pareto optimization problem, improved agreement with the experimental data comes at the expense of growing inconsistencies with the theoretical reference model. Here, we…