18 papers · ranked by Valyu relevance
Shravan Vasishth, Bruno Nicenboim
We present the fundamental ideas underlying statistical hypothesis testing using the frequentist framework. We begin with a simple example that builds up the one-sample t-test from the beginning, explaining important concepts such as the sampling distribution of the sample mean, and the iid assumption. Then we examine…
Jian Gao
Background In medical research and practice, the p-value is arguably the most often used statistic and yet it is widely misconstrued as the probability of the type I error, which comes with serious consequences. This misunderstanding can greatly affect the reproducibility in research, treatment selection in medical…
Mark Rubin
The inflation of Type I error rates is thought to be one of the causes of the replication crisis. Questionable research practices such as p-hacking are thought to inflate Type I error rates above their nominal level, leading to unexpectedly high levels of false positives in the literature and, consequently…
Vance W. Berger, J. Rosser Matthews
It is human nature to try to recognize patterns and to make sense of that which we observe. Unfortunately, our intuition is often wrong, and so there is a need to impose some objectivity on the methods by which observations are converted into knowledge. One definition of biostatistics could be precisely this, the…
Stephen W Marshall
Background Power for assessing interactions during data analysis is often poor in epidemiologic studies. This is because epidemiologic studies are frequently powered primarily to assess main effects only. In light of this, some investigators raise the Type I error rate, thereby increasing power, when testing…
Ruihuan Shen
The current paper is a commentary on the Metabolic reprogramming-associated genes predict overall survival for rectal cancer (Jian-Qing Lin et al 2020). The authors concluded that ‘Patients with high-risk demonstrated significantly poorer survival outcomes than patients with low-risk in the TCGA database. Also…
Joseph F. Mudge, Leanne F. Baker, Christopher B. Edge, Jeff E. Houlahan + 1 more
'Jeff E. Houlahan' 'Zheng Su'] Null hypothesis significance testing has been under attack in recent years, partly owing to the arbitrary nature of setting α (the decision-making threshold and probability of Type I error) at a constant value, usually 0.05. If the goal of null hypothesis testing is to present conclusions…
E. Chasseloup, A. Tessier, M.O. Karlsson
Pharmacometric approaches achieves higher power to detect a drug effect compared to traditional statistical hypothesis tests. Known drawbacks come from the model building process where multiple testing and model misspecification are major causes for type I error inflation. IMA is a new approach using mixture models and…
Andrew P Grieve
Extensions Authors: ['Andrew P Grieve'] In clinical studies upon which decisions are based there are two types of errors that can be made: a type I error arises when the decision is taken to declare a positive outcome when the truth is in fact negative, and a type II error arises when the decision is taken to declare a…
Denes Szucs, John PA Ioannidis
Null hypothesis significance testing (NHST) has several shortcomings that are likely contributing factors behind the widely debated replication crisis of psychology, cognitive neuroscience and biomedical science in general. We review these shortcomings and suggest that, after about 60 years of negative experience, NHST…
Chang‐Xing Ma, Kejia Wang
Measurements are generally collected as unilateral or bilateral data in clinical trials or observational studies. For example, in ophthalmologic studies, statistical tests are often based on one or two eyes of an individual. For bilateral data, recent literatures have shown some testing procedures that take into…
Mark Rubin
During multiple testing, researchers often adjust their alpha level to control the familywise error rate for a statistical inference about a joint union alternative hypothesis (e.g., "H1,1 or H1,2"). However, in some cases, they do not make this inference. Instead, they make separate inferences about each of the…
Rand R. Wilcox, Guillaume A. Rousselet
There is a vast array of new and improved methods for comparing groups and studying associations that offer the potential for substantially increasing power, providing improved control over the probability of a Type I error, and yielding a deeper and more nuanced understanding of neuroscience data. These new techniques…
Moritz Fabian Danzer, Jannik Feld, Andreas Faldum, Rene Schmidt + 1 more
The one-sample log-rank test is the method of choice for single-arm Phase II trials with time-to-event endpoint. It allows to compare the survival of patients to a reference survival curve that typically represents the expected survival under standard of care. The one-sample log-rank test, however, assumes that the…
Daniel J. Schad, Shravan Vasishth
When researchers carry out a null hypothesis significance test, it is tempting to assume that a statistically significant result lowers Prob(H0), the probability of the null hypothesis being true. Technically, such a statement is meaningless for various reasons: e.g., the null hypothesis does not have a probability…
Richard J. Wang, Predrag Radivojac, Matthew W. Hahn
Errors in genotype calling can have perverse effects on genetic analyses, confounding association studies and obscuring rare variants. Analyses now routinely incorporate error rates to control for spurious findings. However, reliable estimates of the error rate can be difficult to obtain because of their variance…
Linde Schoenmaker, Olivier Béquignon, Willem Jespers, Gerard van Westen
Generative deep learning models have emerged as a powerful approach for de novo drug design, as they aid researchers in finding new molecules with desired properties. Despite continuous improvements in the field, a subset of the outputs that sequence-based de novo generators produce cannot be progressed due to errors.…
Authors not listed
This study presents a validation and refinement of the “yellow cards” error detection workflow that can be applied to any property connected to molecular structure. In our implementation the workflow employed 5 predictive models with each assigning a “yellow card” to 5% of the entries with worst prediction accuracy.…