Every start year gets a turn
Monte Carlo asks how often a plan would work if returns were random around some average. Backtesting asks something more concrete and, to many people, more convincing: did this plan survive the past? The test starts your plan in 1900 and runs it to the end of the data. Then 1901. Then 1902, and so on, wrapping around when the series runs out so that every start year gets a full-length retirement. Each start year is a cohort, a person who happened to begin at that moment, and the result is the share of cohorts whose money outlasted them.
History earns its keep here because it is not random. It clusters. The 1970s delivered a decade of inflation and poor real returns in a row (the UK's own price indices from that era tell the same story); 1929 and 2000 both began long recoveries. Random draws rarely produce runs that punishing, so history's worst sequences are nastier than anything a simulation builds around your own assumed return. Yet the two tests can still disagree in either direction, and the reason is worth understanding: the Monte Carlo run uses the return you assumed, while the backtest uses the record, and the American century in this dataset ran richer than a cautious 5% real assumption. Assume modestly and the backtest will often come out kinder overall, even though its failures are uglier. The gap between the two answers is telling you how much of your plan rests on which return story you believe.
Know what the data is before leaning on any of it. The series is an approximation of US real equity returns in the tradition of the Shiller dataset, the same body of evidence behind the Trinity study and the 4% rule it produced. A UK, European or Japanese investor's real history differed, in Japan's case dramatically, and UK data is not currently modelled; this page says so rather than implying a precision it does not have. The withdrawal rate explorer discusses how that affects a sensible rate outside the US. The wrap-around sequences are artificial too, since later start years borrow early data to complete a full retirement, mixing eras that never actually followed one another. And roughly a hundred overlapping cohorts is a small sample that shares data heavily, so a 95% success rate does not mean ninety-five independent successes. Treat the output as shape and intuition, not precision. Above all, the past is not a distribution of the future: the last century contained one country's unusually strong run. Evidence, not a guarantee.
1966, the quiet killer
A handful of start years do nearly all the damage in a long backtest. 1929 and 1930 retired straight into the Depression; 2000 into the dot-com bust. And then there is 1966: no dramatic single-year crash at all, just fifteen years of inflation grinding real returns to nothing while withdrawals continued.
That last one is the important lesson. Most people picture the risk as a spectacular crash, but historically the plans that failed were more often killed by a long, dull stretch of mediocre real returns, much harder to notice while it is happening and much harder to react to. Lengthen the retirement and the effect sharpens: a longer horizon promotes the slow-grind cohorts like 1966 above the dramatic crash years, because sustained mediocre returns hurt more the longer you are exposed to them. Extend your own horizon in the tool and watch that reshuffle happen.
Elena against the century
Elena is 47 with £520,000, saving £900 a month, planning £31,000 a year at 4%. The backtest passes 94% of its cohorts; eight fail, and they start in the early 1960s, at the turn of the century, and in 1998. Notice those are not the classic crash years, and the reason teaches something: Elena spends her first few years still accumulating, so a cohort that starts in 1962 retires into the 1966 grind, and a 1998 start retires into the 2000s crashes. The dangerous years have not changed; her start dates are simply offset from them by the length of her run-up.
Her Monte Carlo run, at around 65%, comes out far harsher, and that gap is the section above in action: her cautious 5% assumption sits well below what the American record actually paid. Both numbers are honest answers to different questions.
Lowering her withdrawal rate to 3.5% clears every cohort in the dataset. A £2,000 spending cut at the same 4% rate, by contrast, only nudges the pass rate to 94.4%, because at these margins the rate moves the target far more than the withdrawals. Insurance here is bought with the rate.
Her result carries the standard caveats: fixed withdrawals, no tax, no fees. Flexible spending changes the picture materially, and withdrawal strategies shows by how much.
So which deserves more trust, this or Monte Carlo? Neither alone. Simulation covers futures history never produced; history covers patterns randomness rarely generates. A plan that passes both comfortably is in genuinely good shape. One that passes only the friendlier of the two is not.