11 papers · ranked by Valyu relevance
Berna Devezer, Erkan O. Buzbas
Replication studies estimate the replicability rate of scientific results by aggregating binary verdicts of experiments. Exact replications are rarely attainable, so most replication sequences are non-exact. Experiments differ in ways that matter and do not share a single data-generating process. We formalize two…
Peng Wang, Hongyuan Cao, Xiaoquan Wen, Macha Nikolski
We further extend (0-3)$(1)$ to model observed effects, along with the corresponding standard errors. The resulting hierarchical model, incorporating a user-defined $P_{\text{mis}}$ threshold, fully characterizes the properties of replicable signals given a group of estimated effects and their corresponding standard…
María Paula Fernández-García, Guillermo Vallejo-Seco, Pablo Livácic-Rojas
The replicability crisis in the behavioral sciences should no longer be understood as a merely technical problem confined to the failure to reproduce specific findings. Rather, it reflects a deeper structural issue: a growing misalignment between data, method, and inference. When these three levels cease to be…
Carole J. Lee
Crises in peer review capacity, study replication, and AI-fabricated science have intensified interest in automated tools for assessing scientific research. However, the scientific community has a history of decontextualizing and repurposing credibility markers in inapt ways. I caution that AI science evaluation tools…
Angermeir, Florian, Amougou, Maximilian + 14 more
Large Language Models have gained remarkable interest in industry and academia. The increasing interest in LLMs in academia is also reflected in the number of publications on this topic over the last years. For instance, alone 78 of the around 425 publications at ICSE 2024 performed experiments with LLMs. Conducting…
Siamak K. Sorooshyari, Manuel A. Rivas, Robert Tibshirani
Despite being ubiquitous in science, clustering remains a technique whose results are not quantitatively scrutinized via a framework. We present an analysis called evaluating replicability via iterative clustering assignments (ERICA) that is applied to a dataset to determine whether clusters are identified in a…
Rita Banzi, Monika Varga, Yuri Andrei Gelsleichter, Constant Vinatier + 2 more
Evidence-based solutions are needed to help improve reproducibility in research. This Consensus View presents a consensus-based list of core reproducibility items for research that has been developed by a multidisciplinary group interested in research, open science, and reproducibility. The set of minimum requirements…
Jonna Brenninkmeijer, Maarten Derksen, Stephanie Meirmans, Jeannette Pols
Ethnography and other empirical studies of replication played a significant role in the sociology of scientific knowledge (SSK) during the 1970s and 1980s. Collins and other proponents of SSK highlighted that exact replication was impossible, knowledge was often tacit and hard to explicate, and that results were always…
Max Hopkins, Russell Impagliazzo, C. Ye
Replicability, introduced by (Impagliazzo et al. STOC '22), is the notion that algorithms should remain stable under a resampling of their inputs (given access to shared randomness). While a strong and interesting notion of stability, the cost of replicability can be prohibitive: there is no replicable algorithm, for…
Jan Walleczek, Nikolaus von Stillfried, Stefan Schmidt, Marc Wittmann + 4 more
This metascientific project studied the replicability of Bem Experiment 1, which had claimed a precognitive effect, i.e., the ability to successfully guess the outcome of future random events (Bem. J Pers Soc Psychol. 2011;100: 407−25). The use of advanced methodologies-based on the advanced meta-experimental protocol…
Shanda Li, Qiuhong Anna Wei, Jingwu Tang, Valerie Chen + 4 more
Reproducing research results from papers and released code is central to scientific progress. Existing works have introduced benchmarks to evaluate whether LLM agents can assist with reproducibility, but they are difficult to scale due to their reliance on substantial manual effort for data curation and evaluation. We…