21 papers · ranked by Valyu relevance
Timothy Crawley, Arthur G. Palmer III
The ability to make robust inferences about the dynamics of biological macromolecules using NMR spectroscopy depends heavily on the application of appropriate theoretical models for nuclear spin relaxation. Data analysis for NMR laboratory-frame relaxation experiments typically involves selecting one of several…
Reem Salman, Ayman Alzaatreh, Hana Sulieman, Shaimaa Faisal + 1 more
'Mohamed Medhat Gaber'] In the past decade, big data has become increasingly prevalent in a large number of applications. As a result, datasets suffering from noise and redundancy issues have necessitated the use of feature selection across multiple domains. However, a common concern in feature selection is that…
Kevin Debeire, Jakob Runge, Andreas Gerhardus, Veronika Eyring
Learning causal graphs from multivariate time series is an ubiquitous challenge in all application domains dealing with time-dependent systems, such as in Earth sciences, biology, or engineering, to name a few. Recent developments for this causal discovery learning task have shown considerable skill, notably the…
Aki Nikolaidis, Anibal Solon Heinsfeld, Ting Xu, Pierre Bellec + 2 more
Increasing the reproducibility of neuroimaging measurement addresses a central impediment to the clinical impact of human neuroscience. Recent efforts demonstrating variance in functional brain organization within and between individuals shows a need for improving reproducibility of functional parcellations without…
Kęstutis Baltakys, Juho Kanniainen, Frank Emmert-Streib
Multilayer networks are attracting growing attention in many fields, including finance. In this paper, we develop a new tractable procedure for multilayer aggregation based on statistical validation, which we apply to investor networks. Moreover, we propose two other improvements to their analysis: transaction…
Susmita Datta, Vasyl Pihur, Somnath Datta
Background Generally speaking, different classifiers tend to work well for certain types of data and conversely, it is usually not known a priori which algorithm will be optimal in any given classification application. In addition, for most classification problems, selecting the best performing classification algorithm…
Shrabanti Chowdhury, Ru Wang, Qing Yu, Catherine J. Huntoon + 7 more
Directed gene/protein regulatory networks inferred by applying directed acyclic graph (DAG) models to proteogenomic data has been shown effective for detecting causal biomarkers of clinical outcomes. However, there remain unsolved challenges in DAG learning to jointly model clinical outcome variables, which often take…
Peter Bühlmann, Nicolai Meinshausen
Large-scale data analysis poses both statistical and computational problems which need to be addressed simultaneously. A solution is often straightforward if the data are homogeneous: one can use classical ideas of subsampling and mean aggregation to get a computationally efficient solution with acceptable statistical…
Mansoureh Maadia, Uwe Aickelin, Hadi Akbarzadeh Khorshidi
—The main aim in ensemble learning is using multiple classifiers' outputs rather than one classifier output to aggregate them for more accurate classification. Generating an ensemble classifier generally is composed of three steps: selecting the base classifier, applying a sampling strategy to generate different…
M Bourel, Badih Ghattas
We present some new density estimation algorithms obtained by bootstrap aggregation like Bagging. Our algorithms are analyzed and empirically compared to other methods found in the statistical literature, like stacking and boosting for density estimation. We show by extensive simulations that ensemble learning are…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Varun Saravanan, Gordon J. Berman, Samuel J. Sober
A common feature in many neuroscience datasets is the presence of hierarchical data structures, most commonly recording the activity of multiple neurons in multiple animals across multiple trials. Accordingly, the measurements constituting the dataset are not independent, even though the traditional statistical…
Rory M. Crean, Joanna S. G. Slusky, Peter M. Kasson, Shina Caroline Lynn Kamerlin
Simulation datasets of proteins (e.g., those generated by molecular dynamics simulations) are filled with information about how the non-covalent interaction network within a protein regulates the conformation and thus function of said protein. Most proteins contain thousands of non-covalent interactions, with most of…
Authors not listed
Plastic mechanical recycling is the conventional technological step towards circularity. In such aspects, complex mixtures of polyolefin blends are often fed into mechanical recycling systems, resulting in moulded products with uncertain quality. To add to the difficulty of heterogeneous feedstocks, the testing of…
Rahi Jain, Wei Xu
Feature selection (FS) reduces the dimensions of high dimensional data. Among many FS approaches, ensemble-based feature selection (EFS) is one of the commonly used approaches. The rank aggregation (RA) step influences the feature selection of EFS. Currently, the EFS approach relies on using a single RA algorithm to…
Haikady N. Nagaraja, Shane Sanders, Alan D Hutson
The relationship between social choice aggregation rules and non-parametric statistical tests has been established for several cases. An outstanding, general question at this intersection is whether there exists a non-parametric test that is consistent upon aggregation of data sets (not subject to Yule-Simpson…
Belinda Boehm, Christopher McNeill, David Huang
Understanding the solution-phase behaviour of organic semiconducting polymers is important for systematically improving the performance of devices based on solution-processed thin films of these molecules. Conventional polymer theory predicts that polymer conformations become more compact as solvent quality decreases…
Maria Ulan, Welf Löwe, Morgan Ericsson, Anna Wingkvist
Regression uses supervised machine learning to find a model that combines several independent variables to predict a dependent variable based on ground truth (labeled) data, i.e., tuples of independent and dependent variables (labels). Similarly, aggregation also combines several independent variables to a dependent…
Authors not listed
Next Generation Risk Assessment (NGRA) promotes animal-free, exposure-informed, and hypothesis-driven approaches to chemical safety assessment. In silico tools, such as quantitative structure-activity relationship (QSAR) models, are valuable new approach methodologies (NAMs) for use in NGRA. However, the practical…
James I Austerberry, Daniel J Belton
The problem of protein aggregation is widely studied across a number of disciplines, where understanding the behaviour of the protein monomer, and its behaviour with co-solutes is imperative in order to devise solutions to the problem. Here we present a method for measuring the kinetics of protein aggregation based on…
Swarup Subudhi, Ghansham Chandel, Vishal Sivasankar, Siddhartha Das
Magnetic nanoparticles (MNPs) have been extensively used for drug delivery, on-demand material deposition, etc. In this study, we demonstrate the capability to extract MNPs on-demand from a magnetic nanoparticle laden drop (MNLD) (i.e., a drop of stable aqueous dispersion of MNPs) suspended inside a highly viscous…