Search · four archives
Search · four archives
21 papers · ranked by Valyu relevance
Danny Lu, Aalim Weljie, Alexander R. de Leon, Yarrow McConnell + 2 more
'Oliver F. Bathe' 'Karen Kopciuk'] Background Variable selection is frequently carried out during the analysis of many types of high-dimensional data, including those in metabolomics. This study compared the predictive performance of four variable selection methods using stability-based selection, a new secondary…
Taneli Pusa, Juho Rousu, Kai Wang
Multi-omics analysis offers a promising avenue to a better understanding of complex biological phenomena. In particular, untangling the pathophysiology of multifactorial health conditions such as the inflammatory bowel disease (IBD) could benefit from simultaneous consideration of several omics levels. However, taking…
Mahdi Nouraie, Samuel Müller
Stability selection is a widely adopted resampling-based framework for high-dimensional structure estimation and variable selection. However, the concept of 'stability' is often narrowly addressed, primarily through examining selection frequencies, or 'stability paths'. This paper seeks to broaden the use of an…
Andreas Mayr, Benjamin Hofner, Elisabeth Waldmann, Tobias Hepp + 2 more
'Sebastian Meyer' 'Olaf Gefeller'] Statistical boosting algorithms have triggered a lot of research during the last decade. They combine a powerful machine learning approach with classical statistical modelling, offering various practical advantages like automated variable selection and implicit regularization of…
Benjamin Hofner, Luigi Boccuto, Markus Göker
Background Modern biotechnologies often result in high-dimensional data sets with many more variables than observations (n≪p). These data sets pose new challenges to statistical analysis: Variable selection becomes one of the most important tasks in this setting. Similar challenges arise if in modern data sets from…
Yidi Deng, Jiadong Mao, Jarny Choi, Kim-Anh Lê Cao
Inferring reproducible relationships between biological variables remains a challenge in the statistical analysis of omics data. For example, methods that identify statistical associations may lack interpretability or reproducibility. The situation can be greatly improved, however, by introducing the measure of…
Janek Thomas, Tobias Hepp, Andreas Mayr, Bernd Bischl
We present a new variable selection method based on modelbased gradient boosting and randomly permuted variables. Model-based boosting is a tool to fit a statistical model while performing variable selection at the same time. A drawback of the fitting lies in the need of multiple model fits on slightly altered data…
Yonghan Kwon, Kyunghwa Han, Young Joo Suh, Inkyung Jung
Stability selection is a variable selection algorithm based on resampling a dataset. Based on stability selection, we propose weighted stability selection to select variables by weighing them using the area under the receiver operating characteristic curve (AUC) from additional modelling. Through an extensive…
Tino Werner
Contamination can severely distort an estimator unless the estimation procedure is suitably robust. This is a well-known issue and has been addressed in Robust Statistics, however, the relation of contamination and distorted variable selection has been rarely considered in literature. As for variable selection, many…
Md Hasinur Rahaman Khan, Anamika Bhadra, Tamanna Howlader
The instability in the selection of models is a major concern with data sets containing a large number of covariates. We focus on stability selection which is used as a technique to improve variable selection performance for a range of selection methods, based on aggregating the results of applying a selection…
Annika Strömer, Nadja Klein, Christian Staerk, Florian Faschingbauer + 2 more
in distributional copula regression Authors: ['Annika Strömer' 'Nadja Klein' 'Christian Staerk' 'Florian Faschingbauer' 'Hannah Klinkhammer' 'Andreas Mayr'] Structured additive distributional copula regression allows to model the joint distribution of multivariate outcomes by relating all distribution parameters to…
Anyou Wang, Rong Hai
Numerous software have been developed to infer the gene regulatory network, a long-standing key topic in biology and computational biology. Yet the slowness and inaccuracy inherited in current software hampers their application to the increasing massive data. Here, we develop a software, FINET (Fast Inferring NETwork)…
Mahdi Nouraie, Connor Smith, Samuel Müller
Stability selection is a versatile framework for structure estimation and variable selection in high-dimensional setting, primarily grounded in frequentist principles. In this paper, we propose an enhanced methodology that integrates Bayesian analysis to refine the inference of inclusion probabilities within the…
Janek Thomas, Tobias Hepp, Andreas Mayr, Bernd Bischl
We present a new variable selection method based on model-based gradient boosting and randomly permuted variables. Model-based boosting is a tool to fit a statistical model while performing variable selection at the same time. A drawback of the fitting lies in the need of multiple model fits on slightly altered data…
Karan Uppal, Eva K. Lee
Recent studies have shown that the ensemble feature selection approaches are essential for generating robust classifiers. Existing methods for aggregating feature lists from different methods require use of arbitrary thresholds for selecting the top ranked features and do not account for classification accuracy while…
Authors not listed
This paper addresses the challenges in cell line development (CLD), the lengthy and ambiguous clone screening in upstream biopharmaceutical production. Typically, only a small subset of the later stages of CLD data is used for manually selecting lead clones. Addressing this issue, we introduce a multivariate data…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Ivan Lorca-Alonso, Miguel Arenas, Ugo Bastolla
In previous studies, we presented site-specific substitution models of protein evolution based on selection on the folding stability of the native state (Stab-CPE), which predict more realistically the evolutionary variability across protein sites. However, those Stab-CPE present qualitative differences from observed…
Roberta Coletti, J. Orestes Cerdeira, Marcos Raydan, Marta B. Lopes
High-dimensional omics data often contain more variables than observations, which negatively impacts the performance of classical data analysis methods. Dimensionality reduction is typically addressed through variable selection strategies that incorporate a penalty term into the model. While effective for selecting…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Authors not listed
This paper presents the Multi Cell-line Kinetic Model (MCKM), a novel generalised kinetic mechanistic model specifically tailored for Ambr15™ fed-batch cultivations of multiple Chinese Hamster Ovary (CHO) cell lines producing different recombinant monoclonal antibodies (mAbs). Unlike traditional models that requires…