Curriculum·R405 Strategy Development and Validation·about 43 min

Statistical honesty

By the end of this lesson you can

  • Compute the number of trades needed to distinguish your expectancy from zero
  • Report a confidence interval rather than a point estimate, and read what it says
  • Explain what a p-value does and does not mean
  • State the replication rate for published findings, and what that implies for your own

Senior · enrolled learners

This lesson opens with Replicating Anomalies, published 2020.

What happened
Kewei Hou, Chen Xue and Lu Zhang assembled a data library of 452 published asset pricing anomalies and attempted to replicate each one under a consistent, conservative methodology, using NYSE breakpoints and value-weighted returns so that the results were not driven by the smallest and least tradeable stocks. Under that treatment, 65 percent of the 452 anomalies could not clear the conventional single-test hurdle of an absolute t-statistic of 1.96. In the trading frictions category the failure rate was 96 percent. Applying the higher multiple-testing hurdle of 2.78, appropriate given how many hypotheses the literature has examined, raised the overall failure rate to 82.1 percent. The work was published in the Review of Financial Studies in 2020.
The decision point
These were published, peer-reviewed findings by professional researchers with better data, more time and stronger incentives for care than any individual will ever have, and roughly four in five did not survive a consistent replication. Nothing in that is an accusation. It is the base rate for the activity you are about to perform on a smaller sample with a weaker method, and it is the number your own results have to be interpreted against.

What you will be able to answer

  • How do you compute the t-statistic on a record?
  • How many trades does a strong edge need?
  • What does the confidence interval say?
  • What is the replication base rate?

Orientation and Year One are open: anyone can read them without an account. From Year Two onward the lessons are for enrolled learners, because progress through the later years only means anything if it is tracked against a record.

It is free. We do not sell the list and there is nothing to buy at the end of it.

Terms used here

Sources and review

Confidence high·Volatility low·Reviewed 2026-08-07·Owner unassigned

Contested

The interpretation of the replication results is genuinely contested. The authors' treatment of microcaps and their use of value weighting are methodological choices that other researchers have argued are too conservative, and several original authors have responded. Per P6 this lesson uses the study as a base rate for how often findings survive a consistent independent replication and does not assert that the failing anomalies are all spurious.

The t-statistic arithmetic in parts one and two assumes independent trades drawn from a stationary distribution. Real trade outcomes are neither, being clustered by regime, which inflates the apparent statistic. The correction is beyond this course and the direction is worth knowing: the true sample requirement is larger than the figures here.

R403-04 owns expectancy and R-multiples. R405-03 owns multiple testing and the trial count. This lesson owns sample size, intervals and what a p-value means. Keep the splits.