― Uncertainty & ensembles · how well probabilities match reality

Reliability diagrams and sharpness: assessing forecast quality

Reliability diagrams evaluate how well forecast probabilities correspond to observed frequencies. They are a critical tool for assessing the quality and calibration of probabilistic weather forecasts, including ensemble predictions.

8 min readUpdated Verified · google/gemini-2.5-flash-liteLearn
SEE THIS AT YOUR SITE Clonmel · Co. Tipperary
ON THIS PAGE
  1. Forecast probability against observed frequency
  2. Reading the diagonal: perfect reliability
  3. Over- and under-confidence in forecasts
  4. Sharpness versus reliability: a balanced view
  5. Binning and sample-size caveats
  6. Rare events and noisy bins
  7. Using diagrams to retune calibration
  8. Questions
  9. Sources

01Forecast probability against observed frequency

A reliability diagram assesses the calibration of a probabilistic forecast. For a forecast to be reliable, the predicted probability of an event should match the observed frequency of that event. For example, if a forecast states there is a 30% chance of wind speeds exceeding a certain threshold, then over many such forecasts, the event should occur approximately 30% of the time.

To construct a reliability diagram, forecast probabilities are grouped into bins (e.g., 0–10%, 10–20%, ..., 90–100%). For each bin, the diagram plots the average forecast probability within that bin against the observed frequency of the event when the forecast fell into that bin. This comparison reveals whether the forecast system is overconfident, underconfident, or well-calibrated. The process involves collecting a large sample of forecasts and corresponding observations, then sorting them by their forecast probability.

Reliability is distinct from resolution, which measures the forecast's ability to discriminate between events and non-events, and sharpness, which relates to the spread of forecast probabilities. A perfectly reliable forecast might still be unsharp if it always predicts a 50% chance for every event, regardless of actual conditions. The goal is often to achieve both reliability and sharpness, as a sharp but unreliable forecast is misleading, and a reliable but unsharp forecast provides little useful information.

02Reading the diagonal: perfect reliability

A reliability diagram typically plots observed frequency on the y-axis against forecast probability on the x-axis. The ideal scenario for a perfectly calibrated forecast system is represented by the diagonal line from (0,0) to (1,1). If all plotted points fall on this diagonal, it means that for any given forecast probability, the event occurred with precisely that frequency. This line is often referred to as the 'line of perfect reliability'.

For instance, if the forecast system predicted a 60% chance of an event, and the corresponding point on the diagram lies on the diagonal, it indicates that the event actually occurred 60% of the time when such a forecast was made. Deviations from this diagonal reveal systematic biases in the forecast probabilities. A point at (0.6, 0.6) means that when the forecast probability was 60%, the event happened 60% of the time.

The further a point lies from the diagonal, the less reliable the forecast for that probability range. Understanding these deviations helps forecasters and users interpret the meaning of a given probability. For example, if the forecast system consistently predicts a 70% chance of rain, but it only rains 50% of the time in those instances, the forecast is overconfident for that probability range, as the observed frequency (0.5) is lower than the forecast probability (0.7).

03Over- and under-confidence in forecasts

Deviations from the diagonal line indicate either over-confidence or under-confidence in the forecast system. If the plotted curve lies below the diagonal, the forecast system is over-confident. This means that the forecast probabilities are too high relative to the observed frequencies. For example, if the forecast states a 90% chance of strong winds, but strong winds only occur 70% of the time when such a forecast is issued, the system is over-confident for high probabilities. The curve would pass through a point like (0.9, 0.7).

Conversely, if the plotted curve lies above the diagonal, the forecast system is under-confident. In this case, the forecast probabilities are too low compared to the observed frequencies. If the forecast predicts a 20% chance of an event, but the event actually occurs 40% of the time, the system is under-confident for low probabilities. The curve would pass through a point like (0.2, 0.4).

An example: A model predicts a 40% chance of gusts exceeding 30 mph. Over 100 instances where this 40% probability was forecast, the gust limit was exceeded 30 times. The observed frequency (0.30) is below the forecast probability (0.40), indicating over-confidence in this probability range. If the limit was exceeded 50 times, the observed frequency (0.50) would be above the forecast, indicating under-confidence. This distinction is crucial for users making decisions based on probabilistic forecasts.

Exceedance curve Clonmel
CHART LOADINGexceedance_curveReading Clonmel…

The exceedance curve shows the probability of exceeding user-defined thresholds. Its calibration is assessed using reliability diagrams.

04Sharpness versus reliability: a balanced view

While reliability is crucial, it is not the only measure of forecast quality. Sharpness refers to the tendency of a forecast system to issue probabilities near 0% or 100%, rather than clustering around 50%. A sharp forecast provides more definitive information, reducing uncertainty. For example, a forecast of 95% chance of rain is sharper than a 55% chance, assuming both are reliable. A perfectly unsharp but reliable forecast would always predict the climatological frequency of an event.

However, sharpness without reliability can be misleading. A system could be very sharp by always predicting 0% or 100%, but if these predictions are often wrong, the forecast is not useful. Conversely, a perfectly reliable forecast that always predicts 50% (e.g., a climatological average) provides little actionable insight, even if it is perfectly calibrated. The ideal forecast is both reliable and sharp, meaning it accurately reflects the observed frequencies and provides distinct probabilities for different outcomes.

Forecasters often seek a balance. A common approach is to use ensemble forecasts, which naturally produce a range of probabilities, and then apply statistical post-processing techniques to improve both reliability and sharpness. The Brier Score, a common metric for probabilistic forecasts, decomposes into reliability, resolution, and uncertainty components, allowing for a comprehensive evaluation.

05Binning and sample-size caveats

The construction of reliability diagrams involves binning forecast probabilities into discrete intervals. The choice of bin size can influence the appearance and interpretation of the diagram. Too few bins might mask important details, while too many bins can lead to insufficient data points within each bin, resulting in a noisy diagram. Typically, 5 to 10 bins are used, such as 0-10%, 10-20%, and so on.

Sample size is a critical caveat. Reliability diagrams require a large number of forecast-observation pairs to be statistically robust. If the sample size is small, especially for certain probability ranges or for rare events, the observed frequencies within those bins can be highly variable and not truly representative of the forecast system's underlying reliability. This can lead to misleading interpretations of the diagram, with points appearing off the diagonal due to random chance rather than systematic bias.

For example, if only 10 forecasts fall into the 90-100% probability bin, and the event occurs 8 times, the observed frequency is 80%. This single point might suggest over-confidence, but with such a small sample, it is difficult to draw firm conclusions. Adequate sample sizes, often thousands of forecasts, are essential for meaningful reliability analysis.

06Rare events and noisy bins

Reliability diagrams can be particularly challenging to interpret for rare events. By definition, rare events occur infrequently, meaning that the bins corresponding to low forecast probabilities (e.g., 0–10%) will contain a large number of forecasts, while bins for high probabilities (e.g., 90–100%) will have very few, if any, instances where the event actually occurred. This leads to noisy bins with high variability in observed frequencies, especially at the higher probability end.

For example, if a severe storm with a 5% chance of occurrence is forecast, there might be many instances of such a forecast, but the event itself happens very rarely. The observed frequency for the 0–10% bin might be well-behaved. However, if a forecast ever assigns a 90% probability to this rare event, there might be only one or two such instances in the entire dataset. If the event occurs in one of those two instances, the observed frequency for that bin would be 50%, which is a highly unstable estimate.

Specialised techniques, such as pooling data over longer periods or using different binning strategies, are sometimes employed to address these issues for rare events. The interpretation must always consider the underlying climatology of the event being forecast.

Observed vs model Clonmel
CHART LOADINGobs_vs_modelReading Clonmel…

This chart compares observed values against model forecasts. Reliability diagrams extend this by comparing probabilistic forecasts against observed frequencies of events.

07Using diagrams to retune calibration

Reliability diagrams are not just diagnostic tools; they are also instrumental in retuning and improving forecast calibration. By identifying systematic biases (over- or under-confidence) in specific probability ranges, forecasters can apply statistical post-processing techniques to adjust the raw model probabilities. This process is often called recalibration or bias correction.

For instance, if a diagram consistently shows that a model is over-confident for probabilities between 60% and 80% (i.e., the curve is below the diagonal in that range), a recalibration algorithm might systematically reduce the forecast probabilities in that range. This could involve mapping the raw forecast probabilities to a new set of probabilities that are closer to the observed frequencies. Common methods include isotonic regression or using logistic regression models.

The goal of recalibration is to shift the points on the reliability diagram closer to the diagonal, thereby improving the trustworthiness of the probabilistic forecasts. This ensures that when The Wind Agent displays an exceedance probability, that number genuinely reflects the observed frequency of the event over time. This continuous feedback loop of verification and recalibration is fundamental to maintaining high-quality probabilistic weather predictions.

Questions

What is the main purpose of a reliability diagram?

The main purpose of a reliability diagram is to assess the calibration of a probabilistic forecast. It shows whether the predicted probability of an event matches the observed frequency of that event over a large sample of forecasts. This helps users understand how trustworthy a given probability forecast is.

How can I tell if a forecast is over-confident from a reliability diagram?

If the curve on a reliability diagram lies below the diagonal line, the forecast system is over-confident. This means that for a given forecast probability, the event occurred less frequently than predicted. For example, if the forecast said 80% chance, but the event only happened 60% of the time, the system is over-confident.

What is the difference between reliability and sharpness?

Reliability refers to how well forecast probabilities match observed frequencies. Sharpness refers to the tendency of a forecast to issue probabilities close to 0% or 100%, providing more definitive information. A forecast can be reliable but unsharp (e.g., always predicting 50%), or sharp but unreliable (e.g., often wrong with high confidence).

Why are reliability diagrams difficult for rare events?

Reliability diagrams are challenging for rare events because there are typically very few instances of the event occurring, especially at higher forecast probabilities. This leads to small sample sizes in certain bins, making the observed frequencies highly variable and potentially unrepresentative. Interpreting these noisy bins requires caution.

How do forecasters use reliability diagrams to improve forecasts?

Forecasters use reliability diagrams to identify systematic biases in their models. If a diagram shows consistent over- or under-confidence in certain probability ranges, statistical post-processing techniques (recalibration) can be applied to adjust the raw model probabilities. This process aims to shift the curve closer to the diagonal, improving the overall calibration and trustworthiness of the forecasts.

SOURCES

  1. WMO Manual on the Global Data-Processing and Forecasting System
  2. ECMWF Forecast Verification
  3. NOAA Ensemble Verification
  4. Met Éireann Weather Forecast Verification

Thresholds on this page are commonly cited figures, attributed to their source — never statutory limits. Modelled forecasts are planning support, not on-site measurement.