Infrastructure projects do not overrun at random. They overrun in one direction, for reasons that repeat — which is what makes the error correctable rather than merely regrettable. The correction is called the outside view, and the data it needs is probably sitting in your own finished jobs.
Across roads, rail, tunnels and fixed links, studied worldwide and over decades, the estimate that sanctioned the project sits materially below the money the project eventually consumed. That much is unremarkable; estimates are uncertain. What is remarkable is the shape of the error. Genuine uncertainty scatters. This does not. It leans, consistently, in the same direction, which means it is not noise to be absorbed but a bias to be measured and corrected.
Bent Flyvbjerg and colleagues, building on the behavioural work of Daniel Kahneman and Amos Tversky, traced the lean to two forces that push the same way. Optimism bias is the cognitive habit of forecasting from the most favourable version of your own plan — the smoothest delivery, the best available productivity, the ground conditions on the good side of the borehole log. It is a habit, not a deception. Strategic misrepresentation is the deliberate shading of numbers by people whose funding or competitive position depends on a low figure. The two are almost impossible to separate in any given estimate, and the distinction matters less than the consequence they share: the base estimate describes a best case while being presented as an expected one.
The Master Library in TX1:Trinity turns an organisation's rates into a shared, linked, synchronised card — the mechanism that makes your completed work queryable instead of archived.
Start a free 14-day trialThe uncomfortable part of Kahneman and Tversky's planning fallacy is that diligence does not cure it. An estimator who works harder inside an optimistic set of assumptions produces a more precise wrong answer. Detail is not accuracy. A thousand-line build-up assembled from the same hopeful productivity rates is simply a very well-documented best case, and the confidence it generates in the room is the most dangerous thing about it.
A note on vocabulary, because the frameworks do not agree on it. The RES Contingency Guideline and Infrastructure Australia and Treasury policy use optimism bias plainly. The Commonwealth DITRDCA guidance prefers to speak of estimate bias and calibration. The concern is identical: an estimate never tested against real outcomes is an estimate that will tend to be exceeded.
Dan Lovallo and Daniel Kahneman drew the line that organises everything else on this page. The inside view forecasts from the specifics of the case in front of you: this scope, this programme, this risk register, modelled bottom-up. The outside view sets those specifics aside and asks how a class of comparable cases actually turned out. Both are legitimate. Only one of them is immune to the plan's own optimism, and only one of them knows anything about this particular project.
Reference class forecasting is the outside view made operational, in three steps. Identify a class of completed projects similar enough in type, scale and delivery context that their outcomes are informative. Establish the distribution of cost overruns across that class. Then read off the uplift at the percentile corresponding to the confidence you need. The insight underneath is almost rude in its simplicity: the best predictor of how your project will perform is how similar projects actually performed, not how confident you feel about your plan.
A bottom-up model of this project: ranged quantities, contingent risks priced as probability by impact, correlation modelled, reported as an S-curve. Precise, specific and interrogable — and quietly inheriting every optimistic assumption nobody challenged.
The historical overrun distribution of comparable completed projects. Blind to what makes this project unusual, immune to what makes this project's plan optimistic. It answers a different question: how do jobs like this usually go?
When the modelled result and the reference-class benchmark land in similar territory, the budget is defended from two independent directions. When they diverge, the gap is a finding to chase — never a pair of numbers to average.
This is not a fringe technique. Flyvbjerg and COWI's 2004 study for the UK Department for Transport is the direct ancestor of the mandatory optimism-bias uplifts in UK Supplementary Green Book guidance, applied to capital costs at early business-case stages and decaying as design matures. The published bands are illustrative upper bounds rather than answers — roughly 24% for standard buildings, around 44% for standard civil engineering, around 66% for non-standard civil work, and north of 200% for equipment and development projects. Read as a diagnosis rather than a prescription, that ladder states bluntly how far an uncorrected early estimate can sit below outturn when nothing is pushing back.
The RES Contingency Guideline lists reference class forecasting among the probabilistic non-simulation methods, credits Kahneman and Tversky, Lovallo and Kahneman, and Flyvbjerg and COWI, and then does something unusual for a guideline: it sets out exactly where the method breaks. These seven are why it stays a cross-check.
| # | Limitation | What it means in practice |
|---|---|---|
| 1 | Assumes the future mirrors the past | Where delivery models, methods or materials have genuinely changed, the historical class stops being representative of the work you are about to do. |
| 2 | Discourages improvement | Uplifting to the historical average entrenches the historical average. If the past overran by a third, budgeting for a third removes the pressure to beat it. |
| 3 | Most organisations lack the data | A credible class needs a clean, comparable, well-populated record of completed projects. Few organisations hold one, and a thin class produces a fragile forecast. |
| 4 | Needs a genuinely similar class | Comparability in type, scale and context is the whole method. A loosely assembled class imports irrelevant outcomes and flatters or punishes the estimate arbitrarily. |
| 5 | Ignores project-specific events | A class-wide uplift cannot see the discrete exposures unique to this job — a particular latent condition, a known approval risk, one difficult interface. |
| 6 | Speaks only to cost | It addresses cost and time overrun and is silent on safety, quality, environment and reputation, which a full risk process must still handle. |
| 7 | Not transparent at driver level | An uplift read off a distribution is one aggregate number. It cannot be decomposed into the risks driving it, tested in a tornado, or re-run when a single assumption moves. |
The one sentence: the outside view needs a reference class, and the reference class most estimators can actually defend is not a published table of uplifts — it is their own completed jobs, if those jobs were ever recorded in a form that can be queried.
Of the seven, the third is the one that actually stops people. Estimators read about the outside view, agree with it, and then discover they have no reference class — no tidy database of comparable projects with sanctioned estimates and final outturns sitting ready to be interrogated. So the idea is filed under things large agencies do, and the estimate goes out uncorrected.
But the premise deserves a harder look. An established contractor or consultancy has delivered dozens of jobs. Every one has a priced estimate, measured quantities, actual rates achieved, and an outcome someone reconciled at final account. That is not an absence of data. It is data in a form nobody can query — scattered across workbooks on individual machines, coded differently on every job because whoever set up the file invented a structure that Friday, with no way to answer a question as simple as what a cubic metre of structural concrete cost the last five times we did it.
And in one important respect your own history beats a published class. A national dataset of rail projects tells you about rail projects. Your own jobs tell you about your crews, your plant fleet, your subcontractors, your clients' approval behaviour, the ground in the region you work in, and the market you buy in. That is a far tighter class on every dimension that limitation four cares about. The obstacle was never that the data did not exist. It was that nothing in the toolchain ever turned it into a structure — and the fix is unglamorous: consistent resource codes, consistent measurement and code sets, and one shared library that every estimate is drawn from rather than a folder of forks.
Having argued for it, the honest qualifications. A back catalogue of your own work is a selected sample, not a random one. You hold the jobs you won, which means you hold the estimates that were competitive, which is a different population from all the estimates you produced. Tenders lost on price are often the ones that were priced correctly, and they are exactly the records most likely to have been deleted.
Scope movement hides too. An overrun recorded as an approved variation does not look like an overrun in the accounts; it looks like a larger contract delivered on budget. Unless variations are separated from cost growth on the original scope, a self-built reference class will understate the very bias it was assembled to expose. And any comparison across years needs normalising for escalation, location and market conditions before the numbers mean anything, with the basis of that normalisation written down.
The framework positioning is worth restating, because it is easy to overreach. RES is explicit that reference class forecasting is more properly a validation or benchmarking practice for quality management and governance, not a contingency determination method, and recommends it complement rather than replace first principles risk analysis supported by quantitative schedule risk analysis. TMR cites Flyvbjerg under project delay rather than as a contingency engine, and the Commonwealth cost breakdown template carries a reference-class field beside the modelled result, not instead of it. That consensus is not a technicality: a funder challenging your contingency wants a model they can interrogate, not a percentile drawn from someone else's history.
Three practical rules follow. Do not add an uplift to a modelled contingency — they are two routes to the same destination and stacking them double counts. Do not average a divergence between the two; investigate it, because a modelled P90 sitting well below the class usually means narrow ranges, missing contingent risks or unmodelled correlation. And let the uplift decay as design matures, because its whole justification is the absence of project-specific information, and that absence is temporary.
The Master Library is TX1:Trinity's answer to the folder of forks: an organisation-wide rate card that projects link to rather than copy. A project draws its resources and built-up rates from the shared library, and the link is live — when a rate is updated centrally, linked projects can take the change; when a project departs from the library, the departure is visible as a deviation instead of vanishing into a private copy. That property is what makes a back catalogue comparable. Fifty estimates built from linked library resources are fifty observations of the same rate; fifty forks are fifty anecdotes.
Because built-up resources recompute as the sum of their contributing component rows, a change to one component cascades to the package rate, to every item using it, and to all allocations. Synchronisation is therefore not a bulk find-and-replace over old files; it is the rate graph re-evaluating. And the portable .tx1 package moves a library between machines and organisations intact, compressed and checksum-validated, which is how a library gets shared without being flattened — the subject of building a library.
The second half is benchmark bands. A reference class answers a project-level question, but most of an estimator's day is spent one level down, deciding whether a particular unit rate is defensible. A band is a low, medium and high range for an item type and unit, authored independently of the build-up it will judge, so a finished rate reports its own position rather than waiting for someone to notice. It does not tell you a rate is wrong. It tells you a rate is unusual, which is the more useful signal, because unusual rates are frequently correct — night work, restricted access, a genuinely difficult site — and the band's job is to force the explanation into the open. It is the out-of-band rate nobody noticed that gets into a tender.
Optimism bias is the well-documented tendency to forecast from the most favourable version of your own plan: the smoothest delivery, the best available productivity, no latent conditions and no rework. It is a cognitive habit rather than dishonesty, and it survives detailed work, because adding detail to an optimistic assumption makes the estimate more precise without making it more accurate.
The inside view forecasts from the specifics of the project in front of you: its scope, its programme, its risk register, modelled bottom-up. The outside view ignores those specifics and asks how a class of similar completed projects actually performed. Lovallo and Kahneman framed the distinction, and the practical point is that each catches what the other misses, so the strongest estimates carry both.
No. The RES Contingency Guideline is explicit that reference class forecasting is more properly a validation and benchmarking practice for quality management and governance, not a contingency determination method, and recommends it complement rather than replace first principles risk analysis supported by quantitative schedule risk analysis. It cannot be interrogated risk by risk, and it cannot be re-run when one assumption changes.
There is no threshold that makes a class valid, because similarity matters more than count. A dozen genuinely comparable jobs in the same market, delivery model and ground conditions will outperform a hundred loosely assembled ones. The practical test is whether you can state what the class has in common and defend excluding the projects you left out.
Not without double counting. The uplift and the modelled contingency are two independent routes to the same destination, so adding them stacks the same uncertainty twice. Use the model to derive the number and the reference class to test whether that number is credible. If the two land in similar territory the budget is defended from two directions; if they do not, the gap is the finding.
Investigate the gap rather than splitting the difference. A modelled P90 sitting well below the reference class usually means the model has inherited the plan's optimism through narrow ranges, missing contingent risks or unmodelled correlation. Averaging the two numbers produces a figure that neither method supports and that nobody can defend at a gate.
A linked Master Library, built-up rates that recompute from their components, and benchmark bands that flag the unusual before a reviewer has to find it.