Abstract:
In the empirical sciences, significance thresholds often determine whether findings are treated as evidence of an effect. This paper studies how likely findings that just meet conventional significance thresholds are to remain significant in replications of the same sample size. To answer this question, we estimate the expected replication probability conditional on a given p-value among published studies for experimental economics, psychology, and social science. We validate this measure by showing it accurately predicts actual replication outcomes, outperforming prediction markets. A finding with a p-value of 0.05 has an expected replication probability ranging from 0.10 to 0.25 across fields. Low replicability reflects small true effects in original studies rather than selective reporting. Significant findings are therefore often informative about the existence of an effect, but frequently overstate its magnitude. Finally, we develop a nonparametric estimator and apply it to economics literatures that use larger samples, finding higher but still low replication probabilities.