Skip to content

P50 and P90 confidence levels, explained

A P50 budget has a 50% chance of being enough; a P90 budget a 90% chance. That much is easy. The harder question is how you produce one from your own priced model rather than quoting a figure off somebody else’s report — and what has to be true of the build-up underneath before a P90 means anything at all.

Start here

What a P-value actually is

A P-value — the P is for percentile — is a cost paired with a probability that it will not be exceeded. On its own a number says nothing useful: “the project will cost $118M” is an assertion with no shelf life. Attach the confidence level — “$118M at P90” — and it becomes a position a funding body can act on, because it states what chance the organisation is taking.

P-values exist because an estimate is a forecast made under uncertainty, not a measurement. Run the cost model thousands of times, drawing a different plausible value for each uncertain input on every pass, and you get a whole distribution of outcomes rather than one figure. P-values name specific points on that distribution so two people can be confident they are discussing the same level of confidence.

P50 is the coin-flip: equally likely to be over or under, the median outcome, the figure a project manager controls to. P90 is the high-confidence figure: only a one-in-ten chance of being exceeded, which is why funders hold it. Intermediate levels appear once the estimate sharpens — P75 at tender award in Queensland, P60 at award in Victoria.

Need a defensible P90 before the gate?

A confidence level is only as good as the model beneath it. TX1:Trinity builds the priced structure a probabilistic run needs — resources, quantities, productivity and margins kept separate and re-runnable, not fused into a lump sum.

Start a free 14-day trial
The master output

Reading the S-curve

Every P-value is a point read off one chart: the cumulative probability curve that a simulation produces. The vertical axis is the probability that cost will not exceed a given value; the horizontal axis is the cost. It runs from nothing-is-enough at the bottom left to any-budget-this-large-is-certainly-enough at the top right, and the stretched S shape comes from outcomes clustering in the middle and thinning into the tails.

Read it across and you get a budget: pick 90% on the vertical axis, trace right to the curve, drop to the cost. Read it up and you get a confidence: trace a candidate allocation up to the curve and across to the axis to see what probability that budget actually buys. The second reading is the one most people never use, and it is the more useful of the two when somebody hands you a fixed envelope.

Cumulative probability S-curve with P50 and P90 marked A cumulative distribution of project cost outcomes. The curve rises from about 4% probability at $80M, through 50% at $100M (the P50), to 90% at $118M (the P90), then flattens towards 100% beyond $140M. The shaded band between $100M and $118M is the risk-based contingency. 0% 25% 50% 75% 90% 100% P50 · $100M 50% CONFIDENCE P90 · $118M 90% CONFIDENCE CONTINGENCY P90 − P50 $80M $100M $118M $140M PROJECT COST OUTCOME → PROBABILITY NOT EXCEEDED

The shape carries an argument of its own. The curve is steepest near the median, where a small increase in budget buys a large increase in confidence, and it flattens past P90, where each additional dollar buys almost nothing. That flattening is why frameworks stop at P90 rather than reaching for P99, and why the gap between P50 and P90 is a real, derived quantity rather than an assumption: on the curve above, a $100M P50 and a $118M P90 imply an $18M contingency that fell out of the analysis and can be re-run in front of a reviewer.

Read across, then down

Choose a confidence on the vertical axis, trace right to the curve, drop to the cost axis. This is how a funding figure is set: the framework names the confidence, the curve names the budget.

Or read up, then across

Trace a candidate allocation up to the curve and across to find the confidence it buys. Use this when a number has been handed down and you need to say what probability of overrun it represents.

Steep middle, flat tail

Confidence is cheap near the median and expensive past P90. The same shape explains why a suspiciously vertical curve is a warning: it means the model was told the project is more certain than it is.

P-value by framework

Which P-value do you fund at?

Every major Australian framework reports both a lower figure and a P90. What differs is which one releases money, where the gap between them is held, and whether confidence steps down once the estimate sharpens at award. Knowing the answer for your funding body is the difference between an estimate that clears a gate and one that gets sent back for re-presentation.

FrameworkLower funding pointAward resetWhat the P90 does
TMR · QueenslandP50 budget, with the project manager holding contingency to P50.Approved Project Delivery Value set at P75 at tender award; the saving returns to the program.Mandated beyond the business-case milestone as total out-turn cost. The P90 minus P50 gap is held at portfolio level.
DITRDCA · CommonwealthValidated P50 out-turn cost, re-escalated at each phase.No award reset; the P50 is revalidated as the estimate matures.Held notionally against the project to represent the funder’s total exposure, released on demonstrated need.
RES · national and statesP50 paired with the Performance Measurement Baseline.Set by organisational risk appetite rather than a fixed reset point.Funds a Management Reserve held at a higher delegation than the baseline.
VictoriaP60 at contract award.Steps down from P90 approval to P60 once the tender price is known.The approval figure at business case, before the market has priced the work.

Read the pattern rather than the rows. No framework funds at a single secret percentile; each reports a lower figure and a P90 and then governs who holds the difference. The genuine differentiator is the award reset, because it is the point at which an organisation admits the estimate has improved and releases held confidence back to the program.

A P-value is a probability of not exceeding a cost, not the cost of a specific scenario. P90 is not the worst case — it is a budget you would beat 9 times out of 10.

The one idea

A P90 is a property of your model, not of your project

Two estimators can price the same job and produce P90s that differ by twenty per cent, not because they disagree about the work but because they built different models. The simulation cannot see the project. It sees the structure you gave it: which values were allowed to vary, how far, and which of them move together. Change that structure and the P90 changes, even though nothing about the bridge has altered.

Which means the defensibility of a confidence level is decided before any simulation runs. Four things have to be true of the build-up underneath. Quantities and rates must be separable, so uncertainty in how much can be ranged independently of uncertainty in how dear. Rates must decompose into resources, because “labour is tight this year” is a statement about a resource, not about a line item. Undefined work must be visible as provisional rather than blended into measured work, because pretending an allowance is a priced quantity hides exactly the uncertainty that should be widest. And the shared drivers must be identifiable, so that the items which will move together can be made to move together.

An estimate that satisfies those four conditions can be ranged honestly by anyone. One that does not will still produce a smooth S-curve and a confident-looking P90, and the number will be worth nothing. This is the difference between a P-value you generated and a P-value you were handed.

Don’t do this

The mean is not the median, and six other ways P-values go wrong

Almost everyone assumes the average cost is the 50/50 figure. On a real project distribution it is not, and the difference lands squarely on a funding decision. Cost distributions are right-skewed: there is a floor on how cheaply work can be built, but a long tail of expensive outcomes — a latent-condition claim, a market spike, a major rework. That tail drags the arithmetic mean above the most likely value while leaving it below the median by probability rank. RES notes that on energy and transport projects the mean, which it calls the Central Estimate, typically sits around P30 to P40.

The consequence is concrete. Fund a project at “the average” believing it is the coin-flip and you have funded somewhere near a P35 — a budget that will be exceeded six or seven times in ten. The only true 50/50 figure is the P50 median, and it is higher than the mean. State all three explicitly — mean, median, P90 — and no reviewer can mistake one for another. The remaining six errors surface just as routinely in independent review.

Treating P90 as the worst case

P90 is beaten nine times in ten, which also means it is exceeded one time in ten. The real worst case lives in the tail beyond it. Presenting P90 as a ceiling invites a governance conversation that cannot end well when the tail arrives.

Adding P50 and P90 contingencies together

They are two points on one curve, not two provisions to stack. The P90 already contains everything in the P50 plus more; the incremental money is P90 minus P50. Adding them double-counts the same risk and produces a figure no framework recognises.

Comparing P-values on different bases

A P90 on the base estimate, a P90 on the project estimate and a P90 on the escalated out-turn are three different numbers. Always state which. Most apparent estimate blow-outs between gates are a real-dollar figure being compared with an out-turn one.

Quoting a P-value with no model behind it

“That is our P90” is not a statement unless it is re-runnable. A defensible figure traces back to ranged line items or risk factors, a dollarised risk register, modelled correlation and a stated iteration count — the machinery covered in the Monte Carlo method.

Chasing P99

Past P90 the curve is nearly flat, so each extra point of confidence costs disproportionately for negligible risk reduction. Frameworks stop at P90 deliberately. A P99 budget is rarely defensible and tends to be read as padding rather than rigour.

A thin P10–P90 spread

If the distance between the P10 and P90 outcomes is implausibly narrow, the model has understated uncertainty — ranges too tight, correlation ignored, or contingent risks missing. DITRDCA explicitly flags narrow spreads with thin tails as a sign of an unrealistic model. A credible P90 needs a credible spread behind it, and a suspiciously precise answer is usually the least trustworthy one in the room.

In TX1:Trinity

Building the model a P90 can stand on

TX1:Trinity does not run the simulation. It produces the thing the simulation needs: a base estimate whose structure is visible all the way down. Rates are built from resources coded by type — labour, plant, material, subcontract, other and composite groups — so when a reviewer asks which rates are exposed to the labour market, the answer is a list rather than an opinion. Built-up resources recompute as the sum of their contributing rows, which means a market-driven change to one component cascades to the package rate, to every item using it and to all allocations, and the sensitivity you would otherwise model by hand can simply be observed.

Quantities stay separate from rates, and both are addressable in the formula engine through project parametrics and line references, so the two kinds of uncertainty a simulation cares about can be ranged independently. The pricing pipeline keeps its layers distinct — direct costs plus overheads, then markup per resource type, then risk, corporate and profit margins — which matters more than it sounds: when contingency is a named margin rather than a rate inflated quietly at source, the P50 base and the risk provision cannot be double-counted.

The rest happens alongside. Running a quantitative risk analysis in a dedicated tool remains a separate exercise, and the source disciplines this page draws on name @RISK as the common one in Australian practice. What TX1:Trinity contributes is a model that can be exported as a portable, checksum-validated .tx1 package and handed to whoever runs the analysis, or to the reviewer who wants to interrogate the estimate that produced the P90 — which, as the estimating process makes plain, is the step where most confidence levels quietly fall apart.

Common questions

Questions that come up at the gate

What is the difference between P50 and P90?

Both are costs paired with a probability of not being exceeded. P50 is the cost with a 50 per cent chance of being enough, so the outcome is equally likely to land above or below it. P90 is the cost with a 90 per cent chance of being enough, exceeded only about one time in ten. They are two points on one curve, not two rival budgets, and the distance between them on the cost axis is the risk-based contingency.

Is P90 the worst case?

No. P90 is a budget you would beat nine times out of ten, which means one run in ten still exceeds it. The genuine worst case sits further out in the tail, past P95 and P99, where the curve has almost flattened. Presenting P90 as a ceiling that cannot be breached misrepresents what the number is and sets an expectation the project cannot honour.

Why is the mean not the same as the P50?

Project cost distributions are right-skewed. There is a floor on how cheaply work can be built, but a long tail of expensive outcomes, and that tail drags the arithmetic average above the most likely value while leaving it below the median by probability rank. RES places the mean, which it calls the Central Estimate, at roughly P30 to P40. Fund at the average believing it is the coin-flip and you have under-budgeted by a full confidence band.

Which P-value does my project have to be funded at?

It depends on the framework. TMR sets the budget at P50 and mandates P90 beyond business case, then resets to an Approved Project Delivery Value at P75 on tender award. Commonwealth DITRDCA approval decisions are made on a validated P50 out-turn with P90 held notionally against the project. RES pairs P50 with the Performance Measurement Baseline and P90 with a Management Reserve at higher delegation. Victoria approves at P90 and resets to P60 at award. All of them report both ends.

What makes a P90 defensible in an independent review?

That it can be reproduced. A reviewer will want the base estimate it was run against, the ranges applied and the reasoning behind them, the dollarised risk register, the correlation assumptions, and the iteration count. If the base estimate is a set of lump sums with nothing underneath, there is nothing to range and the P90 is decoration. A defensible P90 starts as a defensible build-up, and the simulation is the last step, not the first.

Keep reading

Where the curve comes from, and where it sits

Is your P90 built on a model or on a memory?

Price the job from resources up, keep quantities and rates separable, and hand the reviewer something they can re-run.