Search · four archives
Search · four archives
10 papers · ranked by Valyu relevance
Danny Lu, Aalim Weljie, Alexander R. de Leon, Yarrow McConnell + 2 more
'Oliver F. Bathe' 'Karen Kopciuk'] Background Variable selection is frequently carried out during the analysis of many types of high-dimensional data, including those in metabolomics. This study compared the predictive performance of four variable selection methods using stability-based selection, a new secondary…
Taneli Pusa, Juho Rousu, Kai Wang
Multi-omics analysis offers a promising avenue to a better understanding of complex biological phenomena. In particular, untangling the pathophysiology of multifactorial health conditions such as the inflammatory bowel disease (IBD) could benefit from simultaneous consideration of several omics levels. However, taking…
Andreas Mayr, Benjamin Hofner, Elisabeth Waldmann, Tobias Hepp + 2 more
'Sebastian Meyer' 'Olaf Gefeller'] Statistical boosting algorithms have triggered a lot of research during the last decade. They combine a powerful machine learning approach with classical statistical modelling, offering various practical advantages like automated variable selection and implicit regularization of…
Benjamin Hofner, Luigi Boccuto, Markus Göker
Background Modern biotechnologies often result in high-dimensional data sets with many more variables than observations (n≪p). These data sets pose new challenges to statistical analysis: Variable selection becomes one of the most important tasks in this setting. Similar challenges arise if in modern data sets from…
Yonghan Kwon, Kyunghwa Han, Young Joo Suh, Inkyung Jung
Stability selection is a variable selection algorithm based on resampling a dataset. Based on stability selection, we propose weighted stability selection to select variables by weighing them using the area under the receiver operating characteristic curve (AUC) from additional modelling. Through an extensive…
Janek Thomas, Tobias Hepp, Andreas Mayr, Bernd Bischl
We present a new variable selection method based on model-based gradient boosting and randomly permuted variables. Model-based boosting is a tool to fit a statistical model while performing variable selection at the same time. A drawback of the fitting lies in the need of multiple model fits on slightly altered data…
Authors not listed
This paper addresses the challenges in cell line development (CLD), the lengthy and ambiguous clone screening in upstream biopharmaceutical production. Typically, only a small subset of the later stages of CLD data is used for manually selecting lead clones. Addressing this issue, we introduce a multivariate data…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…
Authors not listed
This paper presents the Multi Cell-line Kinetic Model (MCKM), a novel generalised kinetic mechanistic model specifically tailored for Ambr15™ fed-batch cultivations of multiple Chinese Hamster Ovary (CHO) cell lines producing different recombinant monoclonal antibodies (mAbs). Unlike traditional models that requires…