22 papers · ranked by Valyu relevance
Shrabanti Chowdhury, Ru Wang, Qing Yu, Catherine J. Huntoon + 7 more
Background Applying directed acyclic graph (DAG) models to proteogenomic data has been shown effective for detecting causal biomarkers of complex diseases. However, there remain unsolved challenges in DAG learning to jointly model binary clinical outcome variables and continuous biomarker measurements. Results In this…
Kevin Debeire, Jakob Runge, Andreas Gerhardus, Veronika Eyring
Learning causal graphs from multivariate time series is an ubiquitous challenge in all application domains dealing with time-dependent systems, such as in Earth sciences, biology, or engineering, to name a few. Recent developments for this causal discovery learning task have shown considerable skill, notably the…
Khabat Khosravi, Nasrin Attar, Sayed M. Bateni, Changhyun Jun + 5 more
'Dongkyun Kim' 'Mir Jafar Sadegh Safari' 'Salim Heddam' 'Aitazaz Farooque' 'Soroush Abolfathi'] Accurate prediction of daily river flow (Q*t) remains a challenging yet essential task in hydrological modeling, particularly crucial for flood mitigation and water resource management. This study introduces an advanced M5…
Di Zhang
Bootstrapping is a powerful statistical resampling technique for estimating the sampling distribution of an estimator. However, its computational cost becomes prohibitive for large datasets or a high number of resamples. This paper presents a theoretical analysis and design of parallel bootstrapping algorithms using…
Ahmad Talafha
Sparse functional data frequently arise in real-world applications, posing significant challenges for accurate classification. To address this, we propose a novel classification method that integrates functional principal component analysis (FPCA) with Bayesian aggregation. Unlike traditional ensemble methods, our…
Sudhir Kumar, Koichiro Tamura, Sudip Sharma
Long runtimes, high memory demands, and reliance on high-performance computing impede phylogenomic analyses. We review a scalable phylogenomic subsampling with upsampling (PSU) framework, in which small subsamples of sites from a concatenated alignment are expanded by upsampling before inference, and the resulting…
Gabriel Madirolas, Regina Zaghi-Lara, Adam Matic, Alex Gomez-Marin + 1 more
Wisdom of the Crowd is the aggregation of many individual estimates to obtain a better collective one. This effect has an enormous potential from the social point of view, as it means that a decision may be taken more effectively by vote among a large crowd than by a small minority of experts. Wisdom of the Crowd has…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Wancen Mu, Eric Davis, Stuart Lee, Mikhail Dozmorov + 2 more
bootRanges provides fast functions for generation of bootstrapped genomic ranges representing the null sets in enrichment analysis. We show that shuffling or permutation schemes may result in overly narrow test statistics null distributions, while creating new ranges sets with a block bootstrap preserves local genomic…
Rory M. Crean, Joanna S. G. Slusky, Peter M. Kasson, Shina Caroline Lynn Kamerlin
Simulation datasets of proteins (e.g., those generated by molecular dynamics simulations) are filled with information about how the non-covalent interaction network within a protein regulates the conformation and thus function of said protein. Most proteins contain thousands of non-covalent interactions, with most of…
Authors not listed
Plastic mechanical recycling is the conventional technological step towards circularity. In such aspects, complex mixtures of polyolefin blends are often fed into mechanical recycling systems, resulting in moulded products with uncertain quality. To add to the difficulty of heterogeneous feedstocks, the testing of…
Ana Helena Tavares, Ana Silva, Tiago Freitas, Maria Costa + 3 more
Despite the advances on data analysis methodologies in the last decades, most of the traditional regression methods cannot be directly applied to large-scale data. Although aggregation methods are especially designed to deal with large-scale data, their performance may be strongly reduced in ill-conditioned problems…
Mathieu Berthe, Pierre Druilhet, Stéphanie Léger
We consider the problem of model building for rare events prediction in longitudinal follow-up studies. In this paper, we compare several resampling methods to improve standard regression models on a real life example. We evaluate the effect of the sampling rate on the predictive performances of the models. To evaluate…
Rahi Jain, Wei Xu
Feature selection (FS) reduces the dimensions of high dimensional data. Among many FS approaches, ensemble-based feature selection (EFS) is one of the commonly used approaches. The rank aggregation (RA) step influences the feature selection of EFS. Currently, the EFS approach relies on using a single RA algorithm to…
Frédéric Lemoine, Olivier Gascuel
Felsenstein’s bootstrap is the most commonly used method to measure branch support in phylogenetics. Current sequencing technologies can result in massive sampling of taxa (e.g. SARS-CoV-2). In this case, the sequences are very close, the trees are short, and the branches correspond to a small number of mutations…
Maria Ulan, Welf Löwe, Morgan Ericsson, Anna Wingkvist
Regression uses supervised machine learning to find a model that combines several independent variables to predict a dependent variable based on ground truth (labeled) data, i.e., tuples of independent and dependent variables (labels). Similarly, aggregation also combines several independent variables to a dependent…
Annette Spooner, Gelareh Mohammadi, Perminder S. Sachdev, Henry Brodaty + 1 more
'Henry Brodaty' 'Arcot Sowmya' ''] Background Feature selection is often used to identify the important features in a dataset but can produce unstable results when applied to high-dimensional data. The stability of feature selection can be improved with the use of feature selection ensembles, which aggregate the…
Michał Bałchanowski, Urszula Boryczka, Jiayi Ma
The aim of a recommender system is to suggest to the user certain products or services that most likely will interest them. Within the context of personalized recommender systems, a number of algorithms have been suggested to generate a ranking of items tailored to individual user preferences. However, these algorithms…
Kaiwen Wang, Yuqiu Yang, Yusen Xia, Guanghua Xiao + 2 more
With the rise of large-scale genomic studies, large gene lists targeting important diseases are increasingly common. While evaluating each study individually gives valuable insights on specific samples and study designs, the wealth of available evidence in the literature calls for robust and efficient meta-analytic…
Belinda Boehm, Christopher McNeill, David Huang
Understanding the solution-phase behaviour of organic semiconducting polymers is important for systematically improving the performance of devices based on solution-processed thin films of these molecules. Conventional polymer theory predicts that polymer conformations become more compact as solvent quality decreases…
Authors not listed
Next Generation Risk Assessment (NGRA) promotes animal-free, exposure-informed, and hypothesis-driven approaches to chemical safety assessment. In silico tools, such as quantitative structure-activity relationship (QSAR) models, are valuable new approach methodologies (NAMs) for use in NGRA. However, the practical…
František Zapletal, Miroslav Hudec, Miloš Švaňa, Radek Němec
Valuable information for decision-making can be obtained by collecting and analyzing opinions from diverse stakeholder or respondent groups, which usually have different backgrounds and are variously affected by the topics under survey. For this to succeed, it is necessary to manage the uncertainty of respondents’…