25 papers · ranked by Valyu relevance
Lorenzo Trippa, Levi Waldron, Curtis Huttenhower, Giovanni Parmigiani
'Giovanni Parmigiani'] > We consider comparisons of statistical learning algorithms using multiple data sets, via leave-one-in cross-study validation: each of the algorithms is trained on one data set; the resulting model is then validated on each remaining data set. This poses two statistical challenges that need to…
Yuqing Zhang, Christoph Bernau, Giovanni Parmigiani, Levi Waldron
Cross-study validation (CSV) of prediction models is an alternative to traditional cross-validation (CV) in domains where multiple comparable datasets are available. Although many studies have noted potential sources of heterogeneity in genomic studies, to our knowledge none have system atically investigated their…
Atesh Koul, Cristina Becchio, Andrea Cavallo
The ability to replicate a scientific discovery or finding is one of the features that distinguishes science from non-science and pseudo-science (Dunlap, ; Popper, ; Collins, ). In the words of Popper (): “only by [such] repetitions can we convince ourselves that we are not dealing with a mere isolated ‘coincidence,'…
Brian H. Willis, Richard D. Riley
An important question for clinicians appraising a meta-analysis is: are the findings likely to be valid in their own practice-does the reported effect accurately represent the effect that would occur in their own clinical population? To this end we advance the concept of statistical validity-where the parameter being…
Bradley Malin, Khaled El Emam, Urjoshi Sinha, Silvia Figini + 2 more
Cross-validation remains a popular means of developing and validating artificial intelligence for health care. Numerous subtypes of cross-validation exist. Although tutorials on this validation strategy have been published and some with applied examples, we present here a practical tutorial comparing multiple forms of…
Jing Lei
Cross-validation is one of the most popular model selection methods in statistics and machine learning. Despite its wide applicability, traditional cross-validation methods tend to select overfitting models, due to the ignorance of the uncertainty in the testing sample. We develop a new, statistically principled…
Nikolaus Kriegeskorte
Crossvalidation is a method for estimating predictive performance and adjudicating between multiple models. On each of k folds of the process, k-1 of k independent subsets of the data (training set) are used to fit the parameters of each model and the left-out subset (test set) is used to estimate predictive…
Max A Little, Gael Varoquaux, Sohrab Saeb, Luca Lonini + 3 more
'Arun Jayaraman' 'David C Mohr' 'Konrad P Kording'] Title: Abstract This three-part review takes a detailed look at the complexities of cross-validation, fostered by the peer review of Saeb et al.’s paper entitled “The need to approximate the use-case in clinical machine learning.” It contains perspectives by reviewers…
Angela Lopez-del Rio, Alfons Nonell-Canals, David Vidal, Alexandre Perera-Lluna
Binding prediction between targets and drug-like compounds through Deep Neural Networks have generated promising results in recent years, outperforming traditional machine learning-based methods. However, the generalization capability of these classification models is still an issue to be addressed. In this work, we…
Authors not listed
Accurate solubility prediction in supercritical carbon dioxide (scCO2) is crucial for optimizing experimental design by eliminating unnecessary and costly trials at an early stage, thereby streamlining the workflow. A comprehensive solubility database containing 31975 records has been compiled, providing a foundation…
Tianchu Zeng, Hetu Li, Shaoshi Zhang, Yan Quan Tan + 40 more
Machine learning is accelerating biomedical research. Cross-validation is widely used to compare predictive performance - not only to benchmark algorithms, but also to inform scientific applications, such as ranking biomarkers. However, prediction performance estimates across cross-validation folds are not independent.…
Prashanth Athri, Vidhya Murali, Pradyumna Y Muralidhar, Cassandra Königs + 4 more
- 1. Department of Computer Science and Engineering, Amrita School of Engineering, Amrita Vishwa Vidyapeetham, Bengaluru, India - 2. PES Center for Pattern Recognition, Department of Computer Science and Engineering, PES University, Bengaluru, India - 3. Bioinformatics and Medical Informatics, Bielefeld University…
Davide Falessi, Jacky Huang, Likhita Narayana, Jennifer Fong Thai + 1 more
'Burak Turhan'] Abstract. [Context] We are in the shoes of a practitioner who uses previous project releases' data to predict which classes of the current release are defect-prone. In this scenario, the practitioner would like to use the most accurate classifier among the many available ones. A validation technique…
Yanzhao Qian, Dinghao Wang, Qi Xuan Ding, Matthew Greenberg + 1 more
Cross-validation (CV) is a widely used technique in statistical learning for model evaluation and selection. Meanwhile, various of statistical learning methods, such as Generalized Least Square (GLS), Linear Mixed-Effects Models (LMM), and regularization methods are commonly used in genomic predictions, a field that…
Satoshi Usami, Naoya Todo, Kou Murayama
Longitudinal designs provide a strong inferential basis for uncovering reciprocal effects or causality between variables. For this analytic purpose, a cross-lagged panel model (CLPM) has been widely used in medical research, but the use of the CLPM has recently been criticized in methodological literature because…
Lluna Maria Bru-Luna, Manuel Martí-Vilar, César Merino-Soto, José Livia-Segovia + 2 more
'José Livia-Segovia' 'Juan Garduño-Espinosa' 'Filiberto Toledano-Toledano'] Background The person-centered care (PCC) approach plays a fundamental role in ensuring quality healthcare. The Person-Centered Care Assessment Tool (P-CAT) is one of the shortest and simplest tools currently available for measuring PCC. The…
Yoichi Ii, Shintaro Hiro, Yoshiomi Nakazuru
Background The diagnostic likelihood ratio (DLR) and its utility are well-known in the field of medical diagnostic testing. However, its use has been limited in the context of an outcome validation study. We considered that wider recognition of the utility of DLR would enhance the practices surrounding database…
Emily Colby, Eric Bair
Cross-validation is frequently used for model selection in a variety of applications. However, it is difficult to apply cross-validation to mixed effects models (including nonlinear mixed effects models or NLME models) due to the fact that cross-validation requires "out-of-sample" predictions of the outcome variable…
Tieu-Long Phan, Hoang-Son Lai Le, Gia-Bao Truong, The-Chuong Trinh + 4 more
HIV-1 (Human immunodeficiency virus-1) has been causing severe pandemics by attacking the immune system of its host. Left untreated, it can lead to AIDS (acquired immunodeficiency syndrome), where death is inevitable due to opportunistic diseases. Therefore, discovering new antiviral drugs against HIV-1 is crucial.…
Ryan Cory-Wright, Andrés Gómez
We revisit the problem of ensuring strong test-set performance via cross-validation. Motivated by the generalization theory literature, we propose a nested k-fold crossvalidation scheme that selects hyperparameters by minimizing a weighted sum of the usual cross-validation metric and an empirical model-stability…
Xuan-Truc Dinh Tran, Tieu-Long Phan, Van-Thinh To, Ngoc-Vi Nguyen Tran + 4 more
3D pharmacophore models describe the ligand’s chemical interactions in their bioactive conformation. They offer a simple but sophisticated approach to decipher the chemically encoded ligand information, making them a valuable tool in Drug Design. Our research summarized the key studies for applying 3D pharmacophore…
Raheleh Khorsan, Cindy Crawford
Background. Evidence rankings do not consider equally internal (IV), external (EV), and model validity (MV) for clinical studies including complementary and alternative medicine/integrative medicine (CAM/IM) research. This paper describe this model and offers an EV assessment tool (EVAT(c)) for weighing studies…
Authors not listed
Machine learning holds significant promise for accelerating biomarker discovery in clinical proteomics, yet its real-world impact remains limited by widespread methodological pitfalls and unrealistic expectations. In this perspective, we critically examine the integration of machine learning into clinical proteomics…
Authors not listed
Background: Janus Kinase 2 (JAK2) is a key kinase in cellular signal transduction. Its abnormal activation is closely related to various myeloproliferative neoplasms and inflammatory diseases. Developing selective JAK2 inhibitors is an important direction in drug discovery. Accurate prediction of compound inhibitory…
Victoria T. Hunniford, Agnes Grudniewicz, Dean A. Fergusson, Emma Grigor + 2 more
Multicenter preclinical studies have been suggested as a method to improve reproducibility, generalizability and potential clinical translation of preclinical work. In these studies, multiple independent laboratories collaboratively conduct a research experiment using a shared protocol. The use of a multicenter design…