Search · four archives
Search · four archives
11 papers · ranked by Valyu relevance
Yidi Deng, Jiadong Mao, Jarny Choi, Kim-Anh Lê Cao
Inferring reproducible relationships between biological variables remains a challenge in the statistical analysis of omics data. For example, methods that identify statistical associations may lack interpretability or reproducibility. The situation can be greatly improved, however, by introducing the measure of…
Haibin Guan, Maaike van Gerwen, Seunghee Kim-Schultz, Elena Colicino + 2 more
High-dimensional multi-omics data integration offers novel opportunities to characterize complex biological systems. Even though sampling variability frequently compromises findings, particularly in small cohorts, the reproducibility and generalizability of the derived latent structures are insufficiently evaluated. We…
Alan Aw, Lionel Chentian Jin, Nilah Ioannidis, Yun S. Song
Fine-mapping methods, which aim to identify genetic variants responsible for complex traits following genetic association studies, typically assume that sufficient adjustments for confounding within the association study cohort have been made, e.g., through regressing out the top principal components (i.e.…
Jinle Tang, Zhe Zhang, Jian Zhan, Yaoqi Zhou
High-resolution protein structure determination by experimental techniques is notoriously costly and labor intensive. This problem is mostly solved with arrival of deep-learning-based computational prediction by AlphaFold2 but only for those proteins with enough naturally occurring homologous sequences. Here, we…
Amartya Singh, Hossein Khiabanian
A critical step in the computational analysis of single-cell RNA-sequencing (scRNA-seq) counts data is that of normalization, wherein the goal is to reduce biases introduced due to technical sources that obscure the underlying biological variation of interest. This is typically accomplished by scaling the observed…
Marika Ström, Nicole Wagner, Iryna Kolosenko, Åsa M. Wheelock
The R-workflow ropls-ViPerSNet (R orthogonal projections of latent structures with Variable Permutation Selection and Elastic Net) facilitates variable selection, model optimization and significance testing using permutations of OPLS-DA models, with the scaled loadings (p[corr]) as the main metric of significance…
Suzette N. Palmer, Animesh Mishra, Shuheng Gan, Dajiang Liu + 2 more
Microbiome research has been limited by methodological inconsistencies. Taxonomy-based profiling presents challenges such as data sparsity, variable taxonomic resolution, and the reliance on DNA-based profiling, which provides limited functional insight. Multi-omics integration has emerged as a promising approach to…
Ivan Lorca-Alonso, Miguel Arenas, Ugo Bastolla
In previous studies, we presented site-specific substitution models of protein evolution based on selection on the folding stability of the native state (Stab-CPE), which predict more realistically the evolutionary variability across protein sites. However, those Stab-CPE present qualitative differences from observed…
Roberta Coletti, J. Orestes Cerdeira, Marcos Raydan, Marta B. Lopes
High-dimensional omics data often contain more variables than observations, which negatively impacts the performance of classical data analysis methods. Dimensionality reduction is typically addressed through variable selection strategies that incorporate a penalty term into the model. While effective for selecting…
Y. Sapozhnikov, J.T. Van Leuven, J.S. Patel, C.R. Miller
Protein function depends on the stable folding or binding of peptide chains, and the extent to which an amino acid substitution disrupts this stability can be a strong predictor of the mutation’s impact on the protein’s function and the organism’s fitness. This study seeks to understand this relationship in…
Georg Manthey, Miriam Liedvogel, Birgen Haest, Michael Manthey + 1 more
The ability to select statistical models based on how well they fit an empirical dataset is a central tenet of modern bioscience. How well this works, though, depends on how goodness-of-fit is measured. Likelihood and its derivatives (e.g. AIC) are popular and powerful tools when measuring goodness-of-fit, though…