Search · four archives
Search · four archives
19 papers · ranked by Valyu relevance
Tino Werner
In modern data analysis, sparse model selection becomes inevitable once the number of predictors variables is very high. It is well-known that model selection procedures like the Lasso or Boosting tend to overfit on real data. The celebrated Stability Selection overcomes these weaknesses by aggregating models, based on…
Taneli Pusa, Juho Rousu, Kai Wang
Multi-omics analysis offers a promising avenue to a better understanding of complex biological phenomena. In particular, untangling the pathophysiology of multifactorial health conditions such as the inflammatory bowel disease (IBD) could benefit from simultaneous consideration of several omics levels. However, taking…
Mahdi Nouraie, Samuel Müller
Stability selection is a widely adopted resampling-based framework for high-dimensional structure estimation and variable selection. However, the concept of 'stability' is often narrowly addressed, primarily through examining selection frequencies, or 'stability paths'. This paper seeks to broaden the use of an…
Yidi Deng, Jiadong Mao, Jarny Choi, Kim-Anh Lê Cao
Inferring reproducible relationships between biological variables remains a challenge in the statistical analysis of omics data. For example, methods that identify statistical associations may lack interpretability or reproducibility. The situation can be greatly improved, however, by introducing the measure of…
Yonghan Kwon, Kyunghwa Han, Young Joo Suh, Inkyung Jung
Stability selection is a variable selection algorithm based on resampling a dataset. Based on stability selection, we propose weighted stability selection to select variables by weighing them using the area under the receiver operating characteristic curve (AUC) from additional modelling. Through an extensive…
Xiaozhu Zhang, Jacob Bien, Armeen Taeb
We study the problem of linear feature selection when features are highly correlated. This setting presents two main challenges. First, how should false positives be defined? Intuitively, selecting a null feature that is highly correlated with a true one may be less problematic than selecting a completely uncorrelated…
Annika Strömer, Nadja Klein, Christian Staerk, Florian Faschingbauer + 2 more
in distributional copula regression Authors: ['Annika Strömer' 'Nadja Klein' 'Christian Staerk' 'Florian Faschingbauer' 'Hannah Klinkhammer' 'Andreas Mayr'] Structured additive distributional copula regression allows to model the joint distribution of multivariate outcomes by relating all distribution parameters to…
Mahdi Nouraie, Connor Smith, Samuel Müller
Stability selection is a versatile framework for structure estimation and variable selection in high-dimensional setting, primarily grounded in frequentist principles. In this paper, we propose an enhanced methodology that integrates Bayesian analysis to refine the inference of inclusion probabilities within the…
Maryam Sadiq, Nasser A. Alsadhan, Ramla Shah, Sidra Younas + 2 more
'Zahid Rasheed' 'Suyan Tian'] Variable selection methods are very popular, especially in the field of big data with large predictors. These procedures improve the accuracy and performance of the model by eliminating irrelevant and redundant variables. The main contribution of this study is to couple a logit model with…
Francis Okyere, Michael Nyanney, Zakariya Yahya Algamal
Penalized logistic regression is widely used in biomedical classification to address multicollinearity and improve predictive performance, yet the stability and reproducibility of selected predictors are often overlooked. This study evaluates feature stability and interpretability in ridge, lasso, and elastic-net…
Thomas M. Lange, Mehmet Gültas, Armin O. Schmitt, Felix Heinrich
Machine learning is frequently used to make decisions based on big data. Among these techniques, random forest is particularly prominent. Although random forest is known to have many advantages, one aspect that is often overseen is that it is a non-deterministic method that can produce different models using the same…
Haibin Guan, Maaike van Gerwen, Seunghee Kim-Schultz, Elena Colicino + 2 more
High-dimensional multi-omics data integration offers novel opportunities to characterize complex biological systems. Even though sampling variability frequently compromises findings, particularly in small cohorts, the reproducibility and generalizability of the derived latent structures are insufficiently evaluated. We…
Alan Aw, Lionel Chentian Jin, Nilah Ioannidis, Yun S. Song
Fine-mapping methods, which aim to identify genetic variants responsible for complex traits following genetic association studies, typically assume that sufficient adjustments for confounding within the association study cohort have been made, e.g., through regressing out the top principal components (i.e.…
Authors not listed
This paper addresses the challenges in cell line development (CLD), the lengthy and ambiguous clone screening in upstream biopharmaceutical production. Typically, only a small subset of the later stages of CLD data is used for manually selecting lead clones. Addressing this issue, we introduce a multivariate data…
Suzette N. Palmer, Animesh Mishra, Shuheng Gan, Dajiang Liu + 2 more
Microbiome research has been limited by methodological inconsistencies. Taxonomy-based profiling presents challenges such as data sparsity, variable taxonomic resolution, and the reliance on DNA-based profiling, which provides limited functional insight. Multi-omics integration has emerged as a promising approach to…
Authors not listed
Solubility is critical in drug discovery and development, as it significantly influences a medication's bioavailability and therapeutic efficacy. Understanding solubility at the early stages of drug discovery is essential for minimizing resource consumption and enhancing the likelihood of clinical success via…
Ivan Lorca-Alonso, Miguel Arenas, Ugo Bastolla
In previous studies, we presented site-specific substitution models of protein evolution based on selection on the folding stability of the native state (Stab-CPE), which predict more realistically the evolutionary variability across protein sites. However, those Stab-CPE present qualitative differences from observed…
Roberta Coletti, J. Orestes Cerdeira, Marcos Raydan, Marta B. Lopes
High-dimensional omics data often contain more variables than observations, which negatively impacts the performance of classical data analysis methods. Dimensionality reduction is typically addressed through variable selection strategies that incorporate a penalty term into the model. While effective for selecting…
Kan Hatakeyama-Sato, Seigo Watanabe, Naoki Yamane, Yasuhiko Igarashi + 1 more
Materials informatics and cheminformatics struggle with data scarcity, hindering the extraction of significant relationships between structures and properties. The "Ugly Duckling" theorem, suggesting the difficulty of data processing without assumptions or prior knowledge, exacerbates this problem. Current…