― Uncertainty & Ensembles · ensuring forecasts match reality

Calibration: making probabilities honest

Calibration adjusts raw ensemble forecasts to ensure their stated probabilities accurately reflect observed frequencies. It corrects for systematic biases and under-dispersion, making probabilistic forecasts reliable for decision-making.

9 min readUpdated Verified · google/gemini-2.5-flash-liteLearn
SEE THIS AT YOUR SITE Clonmel · Co. Tipperary
ON THIS PAGE
  1. Raw ensembles are usually under-dispersive
  2. Bias correction against local observations
  3. Spread inflation
  4. Quantile mapping
  5. Training window and sample size
  6. Calibrating per site and sector
  7. When we decline to calibrate
  8. Calibration provenance in the evidence record
  9. Questions
  10. Sources

01Raw ensembles are usually under-dispersive

Numerical weather prediction (NWP) ensembles, such as the European Centre for Medium-Range Weather Forecasts (ECMWF) Ensemble Prediction System (EPS) or the Global Ensemble Forecast System (GEFS), are designed to quantify forecast uncertainty. They do this by running multiple model simulations from slightly perturbed initial conditions and using different model physics.

However, raw ensemble output often exhibits under-dispersion. This means the spread of the ensemble members (the range of forecast outcomes) is typically narrower than the actual observed variability in the atmosphere. If an ensemble states there is a 20% chance of a certain event, but that event actually occurs 40% of the time under those conditions, the ensemble is under-dispersive and its probabilities are not reliable.

Under-dispersion arises from several factors:

  • Incomplete representation of initial condition uncertainty: Perturbations might not fully capture the true range of initial state errors.
  • Model error: The NWP model itself has limitations and simplifications that are not fully accounted for by ensemble perturbations.
  • Resolution: Coarser resolution ensemble models may smooth out small-scale atmospheric features that contribute to real-world variability.

For decision-makers, under-dispersive ensembles can lead to overconfidence in a single forecast outcome or to underestimation of risk. This necessitates calibration to align the forecast probabilities with observed frequencies, making them statistically reliable.

02Bias correction against local observations

A fundamental step in calibration is bias correction. This addresses systematic errors where the model consistently over- or under-predicts a particular variable, such as wind speed. Bias can arise from local topographic effects not fully resolved by the model, or from systematic model physics errors.

For example, a model might consistently forecast wind speeds 2 m/s higher than observed at a particular coastal site due to an inadequate representation of sea-breeze effects or local terrain channelling. To correct this, a mean bias correction can be applied. If the model's average forecast for a given hour and month is 10 m/s, but the observed average is 8 m/s, a simple bias correction would subtract 2 m/s from all future forecasts under similar conditions.

More sophisticated bias correction methods can account for non-linear biases, where the error might vary with the magnitude of the wind speed itself (e.g., a larger over-prediction at higher speeds). This often involves comparing historical model forecasts against a long record of local observations.

The Wind Agent uses local, quality-controlled observations from Met Éireann and other trusted sources for bias correction. This ensures that the modelled wind speeds are adjusted to better reflect the specific conditions of Irish locations.

Observed vs model Clonmel
CHART LOADINGobs_vs_modelReading Clonmel…

This chart shows how model forecasts compare against actual observations at a specific site. A consistent offset indicates bias that calibration aims to correct.

03Spread inflation

After bias correction, the next step often addresses the under-dispersion of the ensemble. Spread inflation aims to widen the ensemble distribution to match the observed variability. This is typically achieved by scaling the ensemble spread by a factor greater than one.

Consider an ensemble where the mean forecast wind speed is 15 m/s, and the standard deviation (a measure of spread) across its members is 2 m/s. If historical data shows that the actual observed variability around a 15 m/s forecast is typically 3 m/s, then the ensemble spread needs to be inflated. The inflation factor would be 3 m/s / 2 m/s = 1.5.

Each ensemble member F_i would then be transformed: F_i' = Mean + InflationFactor * (F_i - Mean). So, if a member originally forecast 17 m/s (2 m/s above the mean), its inflated value would be 15 + 1.5 * (17 - 15) = 15 + 1.5 * 2 = 18 m/s.

This process effectively stretches the ensemble distribution, making it wider and more representative of the true uncertainty. Spread inflation is vital for ensuring that the probabilities derived from the ensemble (e.g., the chance of exceeding a certain threshold) are statistically reliable. Without it, the ensemble would consistently underestimate the likelihood of extreme events.

04Quantile mapping

Quantile mapping, also known as distribution mapping, is a more advanced calibration technique that goes beyond simple bias and spread correction. It aims to transform the entire forecast distribution to match the observed distribution, quantile by quantile.

The process involves:

  1. Empirical Cumulative Distribution Functions (CDFs): Constructing the CDFs for historical raw forecasts and historical observations for a specific location and forecast lead time.
  2. Mapping: For a new raw forecast, its value is mapped to its corresponding quantile in the raw forecast CDF. Then, the value at that same quantile in the observed CDF is taken as the calibrated forecast.

For example, if a raw forecast of 10 m/s corresponds to the 75th percentile in the historical raw forecast distribution, and the 75th percentile in the historical observed distribution is 9 m/s, then the calibrated forecast would be 9 m/s. This method effectively corrects for both bias and higher-order moments of the distribution, such as skewness.

Quantile mapping is particularly effective for non-Gaussian distributions, which are common for wind speed. It ensures that not only the mean and spread but also the frequency of extreme events are well-represented. The Wind Agent employs quantile mapping to refine its probabilistic forecasts, ensuring that the exceedance probabilities displayed by the fan are as accurate as possible.

Exceedance curve Clonmel
CHART LOADINGexceedance_curveReading Clonmel…

The exceedance curve shows the probability of exceeding various wind speed thresholds. Calibration ensures these probabilities are reliable and match observed frequencies.

05Training window and sample size

The effectiveness of any calibration method, especially those based on statistical models like quantile mapping, heavily depends on the training window and the sample size of the historical data used. A training window refers to the period over which historical forecasts and observations are collected to 'train' the calibration model.

  • Length of Training Window: A longer training window (e.g., 5-10 years) generally provides a more robust statistical sample, capturing a wider range of meteorological conditions and reducing the impact of short-term anomalies. However, too long a window might include data from older model versions that have significantly different biases, making the calibration less relevant to current model performance.
  • Sample Size: The number of forecast-observation pairs within the training window is critical. Sufficient data points are needed to accurately estimate the underlying distributions and their relationships. For instance, calibrating for rare extreme events requires a very large sample size to ensure these events are adequately represented in the training data.

The Wind Agent typically uses a rolling training window of several years (e.g., 3-5 years) to balance the need for a large sample with the desire for relevance to the current model's characteristics. This window is continuously updated to incorporate the latest model improvements and observational data. Insufficient data for a specific location or variable may lead to less robust calibration, which is noted in the evidence record.

06Calibrating per site and sector

Wind conditions are highly localised, influenced by topography, land use, and proximity to the coast. Therefore, a 'one-size-fits-all' calibration approach is insufficient. The Wind Agent performs calibration per site and per sector where sufficient observational data is available.

  • Site-specific calibration: This accounts for fixed local effects, such as a mountain range causing sheltering or acceleration, or a city's roughness. A model's bias in Dublin Airport will differ from its bias in Cork Harbour, and calibration adjusts for these unique signatures.
  • Sector-specific calibration: Wind speed and direction biases can vary significantly depending on the wind's origin. For instance, a model might over-predict southerly winds at a site due to a particular terrain feature, while being accurate for westerly winds. By calibrating separately for different wind sectors (e.g., N, NE, E, SE, S, SW, W, NW), the accuracy of the directional forecast can also be improved.

This granular approach ensures that the calibrated probabilities are as relevant as possible to the specific conditions a user faces. For locations without direct observations, calibration parameters are derived from nearby, climatically similar sites, or from regional calibration models, with the associated uncertainty flagged.

07When we decline to calibrate

While calibration significantly enhances the reliability of probabilistic forecasts, there are specific circumstances where The Wind Agent may decline to apply calibration or will explicitly flag its limitations:

  1. Insufficient Observational Data: If a location lacks a sufficiently long or high-quality record of observations, there is no reliable ground truth against which to train the calibration model. Attempting to calibrate with sparse data can introduce more error than it corrects.
  2. Unstable Model Performance: If the underlying NWP model undergoes frequent, significant changes that alter its bias characteristics, a static calibration model may quickly become outdated. In such cases, a more cautious approach is taken, or calibration is temporarily suspended until a stable performance period allows for retraining.
  3. Extremely Complex Terrain: In areas with exceptionally complex microclimates or very steep, unresolved terrain, even advanced calibration methods may struggle to accurately correct model errors. The model's inherent inability to resolve these features at its grid scale can lead to irreducible uncertainty.
  4. Very Short Lead Times: For very short-range forecasts (e.g., less than 1-2 hours), the model's initialisation and rapidly evolving atmospheric state make statistical calibration less impactful compared to nowcasting techniques. The focus shifts more towards direct observation and very high-resolution models.

In these situations, the instrument will default to displaying the raw ensemble output (after any general, non-site-specific bias correction) and will clearly indicate that the probabilities are uncalibrated, urging users to interpret them with increased caution. The Agreement Spine will also show larger discrepancies.

08Calibration provenance in the evidence record

Transparency in meteorological data is paramount. For every forecast provided by The Wind Agent, the evidence record contains detailed provenance of the calibration applied. This allows users to understand the basis of the probabilistic information they are using.

The evidence record for calibration typically includes:

  • Calibration Method: The specific algorithm used (e.g., quantile mapping, bias-correction + spread inflation).
  • Training Data Period: The start and end dates of the observational and forecast data used to train the calibration model.
  • Observational Source: The specific Met Éireann station or other verified source providing the ground truth observations.
  • Model Version: The version of the underlying NWP model for which the calibration was developed.
  • Calibration Status: Whether the forecast is fully calibrated, partially calibrated (e.g., only bias-corrected), or uncalibrated.
  • Performance Metrics: For calibrated forecasts, metrics like the Brier Score or reliability diagrams can be referenced to demonstrate the historical performance of the calibration.

This level of detail enables users, particularly in high-stakes operations, to assess the confidence they can place in the forecast probabilities. If the calibration data is sparse or outdated, the user can factor that into their decision-making process, aligning with The Wind Agent's principle of providing actionable, transparent information.

Questions

Why is calibration necessary for ensemble forecasts?

Calibration is necessary because raw ensemble forecasts are often under-dispersive, meaning their stated probabilities are too narrow and do not accurately reflect the true uncertainty or observed frequencies of events. Calibration adjusts these probabilities to make them statistically reliable, ensuring that a 10% forecast probability corresponds to an event occurring 10% of the time historically.

What is the difference between bias correction and spread inflation?

Bias correction addresses systematic errors where the model consistently over- or under-predicts values, shifting the mean of the forecast distribution to match observations. Spread inflation, on the other hand, widens the forecast distribution (increases its variance) to account for the fact that raw ensembles often underestimate the true range of atmospheric variability.

How does quantile mapping improve calibration?

Quantile mapping is a more comprehensive calibration technique that transforms the entire forecast distribution to match the observed distribution, quantile by quantile. This corrects not only for bias and spread but also for higher-order distributional properties like skewness, providing a more accurate representation of the full range of possible outcomes.

Why is site-specific and sector-specific calibration important?

Wind conditions are highly localised due to terrain, land use, and proximity to water. Site-specific calibration accounts for fixed local effects unique to a location, while sector-specific calibration addresses biases that vary depending on the wind's direction of origin. This granular approach ensures the calibrated forecasts are as relevant as possible to the specific conditions encountered.

What happens if there isn't enough data for calibration?

If there is insufficient high-quality observational data for a location, The Wind Agent may decline to apply full calibration or will explicitly flag the limitations. In such cases, the instrument will typically present raw ensemble output (after any general, non-site-specific bias correction) and advise users to interpret the probabilities with increased caution, as they are not statistically reliable.

SOURCES

  1. ECMWF: Ensemble forecasting
  2. WMO: Guide to Ensemble Prediction Systems
  3. NOAA: Global Ensemble Forecast System (GEFS)
  4. Met Éireann: Weather Observations
  5. Wilks, D. S. (2011). Statistical Methods in the Atmospheric Sciences (3rd ed.). Academic Press.

Thresholds on this page are commonly cited figures, attributed to their source — never statutory limits. Modelled forecasts are planning support, not on-site measurement.