Forecast verification and skill: measuring how good a forecast really is
Forecast verification is the process of objectively assessing the quality of a weather prediction against observations. It uses metrics like bias, MAE, RMSE for speed, and specific measures for direction and probabilistic forecasts. Understanding these metrics is crucial for interpreting forecast confidence.
ON THIS PAGE
01Bias, MAE, and RMSE for wind speed
When evaluating wind speed forecasts, three common metrics are bias, Mean Absolute Error (MAE), and Root Mean Square Error (RMSE). These provide different perspectives on forecast performance.
- Bias measures the average difference between forecast and observed values. A positive bias indicates the forecast tends to be too high, while a negative bias means it tends to be too low. It is calculated as
Bias = mean(forecast - observation). A perfect forecast has a bias of 0, but bias alone does not tell you about the magnitude of individual errors.
- Mean Absolute Error (MAE) calculates the average magnitude of the errors, without regard to their direction. It is less sensitive to large outliers than RMSE.
MAE = mean(|forecast - observation|). If a forecast consistently over-predicts by 2 m/s and under-predicts by 2 m/s, the bias might be near zero, but the MAE would be 2 m/s.
- Root Mean Square Error (RMSE) is similar to MAE but penalises larger errors more heavily because it squares the differences before averaging. This makes it a good measure when large errors are particularly undesirable.
RMSE = sqrt(mean((forecast - observation)^2)). RMSE is always greater than or equal to MAE.
Worked Example: Consider a series of 5 forecasts (F) and observations (O) for wind speed in m/s:
| Forecast (F) | Observation (O) | F - O | F - O | (F - O)² | ||
|---|---|---|---|---|---|---|
| 10 | 8 | 2 | 2 | 4 | ||
| 12 | 13 | -1 | 1 | 1 | ||
| 9 | 10 | -1 | 1 | 1 | ||
| 11 | 9 | 2 | 2 | 4 | ||
| 8 | 7 | 1 | 1 | 1 |
- Bias:
(2 - 1 - 1 + 2 + 1) / 5 = 3 / 5 = 0.6 m/s. The forecast has a slight positive bias. - MAE:
(2 + 1 + 1 + 2 + 1) / 5 = 7 / 5 = 1.4 m/s. - RMSE:
sqrt((4 + 1 + 1 + 4 + 1) / 5) = sqrt(11 / 5) = sqrt(2.2) ≈ 1.48 m/s.
These metrics are typically computed over many forecast instances and locations to establish a robust track record for a model or a specific forecast product.
02Direction error and circular measures
Wind direction is a circular variable, meaning that a 350° forecast and a 10° observation are only 20° apart, not 340°. Standard linear error metrics like MAE or RMSE are inappropriate for direction.
Instead, circular statistics are used. The most common approach is to calculate the smallest angular difference between the forecast and observation. For example, if the forecast is 350° and the observation is 10°, the error is min(|350 - 10|, 360 - |350 - 10|) = min(340, 20) = 20°. This is then averaged over many instances.
Another method involves decomposing the wind vector into its U (zonal, east-west) and V (meridional, north-south) components and computing errors for these. This provides a more physically meaningful error measure, as it accounts for both speed and direction errors simultaneously. For example, a forecast of 10 m/s from 270° (due west) has U=-10, V=0. An observation of 10 m/s from 0° (due north) has U=0, V=10. The vector error would be substantial, reflecting both the direction and component differences.
Directional accuracy is particularly important for operations sensitive to crosswind limits, such as crane operations or aviation. A small error in direction can lead to a significant change in the crosswind component, even if the speed forecast is accurate.
03Hit rate and false alarm rate for thresholds
For many operational decisions, the primary concern is whether a specific wind speed threshold will be exceeded. Verification for such 'yes/no' events uses metrics like hit rate, false alarm rate, and critical success index.
Consider a threshold of 10 m/s. For a series of forecasts and observations, we can construct a contingency table:
| Observed > 10 m/s | Observed ≤ 10 m/s | |
|---|---|---|
| Forecast > 10 m/s | Hits (H) | False Alarms (FA) |
| Forecast ≤ 10 m/s | Misses (M) | Correct Negatives (CN) |
- Hit Rate (Probability of Detection):
H / (H + M). This is the proportion of observed threshold exceedances that were correctly forecast. A high hit rate is desirable, meaning few events are missed.
- False Alarm Rate:
FA / (FA + CN). This is the proportion of non-events that were incorrectly forecast as events. A low false alarm rate is desirable, meaning few unnecessary alerts.
- False Alarm Ratio:
FA / (H + FA). This is the proportion of 'exceedance' forecasts that were incorrect. This is often more useful in practice than the false alarm rate.
- Critical Success Index (CSI):
H / (H + FA + M). This metric balances hits, false alarms, and misses. It ranges from 0 (no skill) to 1 (perfect forecast). CSI is particularly useful for rare events.
These metrics help evaluate the utility of a forecast system for specific operational thresholds. The Wind Agent's exceedance fan (chart id fan) and exceedance curve (chart id exceedance_curve) directly address these probabilities for your chosen limits.
04Brier Score for probabilities
When forecasts provide probabilities of an event (e.g., P(wind speed > 15 m/s) = 70%), the Brier Score (BS) is a widely used metric for verification. It measures the mean squared difference between the forecast probability and the observed outcome (which is 1 if the event occurred, 0 if it did not).
The formula for the Brier Score is: BS = (1/N) * Σ(pᵢ - oᵢ)², where pᵢ is the forecast probability for the i-th instance, oᵢ is the observed outcome (1 or 0), and N is the number of forecast instances.
Properties of the Brier Score:
- It ranges from 0 to 1. A perfect forecast (where
pᵢis 1 whenoᵢis 1, andpᵢis 0 whenoᵢis 0) has a Brier Score of 0. The worst possible score is 1. - It is a proper scoring rule, meaning forecasters are incentivised to state their true beliefs. Any attempt to 'hedge' or misrepresent probabilities will result in a worse (higher) Brier Score over the long run.
- It can be decomposed into three components: reliability, resolution, and uncertainty. Reliability measures how well the forecast probabilities match the observed frequencies. Resolution measures the forecast's ability to discriminate between event and non-event days. Uncertainty is inherent in the observations and cannot be influenced by the forecast.
The Wind Agent's exceedance fan and curve are built directly from ensemble probabilities, which are verified using metrics like the Brier Score. This ensures the probabilities presented are as reliable as possible given the underlying model's performance.
The exceedance curve plots the probability of exceeding various thresholds. The reliability of these probabilities is assessed using metrics like the Brier Score.
05Skill relative to persistence
It is not enough for a forecast to be 'good'; it must also be better than a simple, easily obtainable forecast. This is where the concept of skill comes in. Forecast skill is typically measured relative to a reference forecast, often persistence or climatology.
- Persistence forecast: This assumes that tomorrow's weather will be the same as today's. For wind, this means assuming the wind speed and direction observed now will continue into the future. While simple, persistence can be a surprisingly good forecast for very short lead times (e.g., 1-2 hours) or in very stable weather patterns.
- Climatology forecast: This assumes that tomorrow's weather will be the average weather for that day of the year at that location, based on historical records. Climatology provides a baseline for long-range forecasts, as it has no skill beyond capturing seasonal cycles.
Skill Score: A common way to express skill is: Skill Score = (Score_forecast - Score_reference) / (Score_perfect - Score_reference). For error-based scores like MAE or RMSE, a lower score is better, so the formula might be adapted to Skill Score = (RMSE_reference - RMSE_forecast) / (RMSE_reference - RMSE_perfect). A perfect forecast has a skill score of 1, while a forecast no better than the reference has a skill score of 0. A negative skill score means the forecast is worse than the reference.
For example, if the RMSE of a wind speed forecast is 2 m/s, and the RMSE of a persistence forecast is 4 m/s, and the perfect RMSE is 0, the skill score would be (4 - 2) / (4 - 0) = 2 / 4 = 0.5. This indicates the forecast is 50% better than persistence.
Forecast skill generally decreases with increasing lead time, as the atmosphere's inherent predictability limits become more apparent. The Wind Agent's 'Agreement Spine' (chart id agreement_strip) implicitly shows how models diverge from persistence and climatology over time.
The Agreement Spine shows how different models and observations align or diverge, giving an intuitive sense of forecast skill and uncertainty over lead time.
06Verification by lead time and sector
Forecast skill is not uniform; it varies significantly with lead time and for different meteorological parameters or geographic sectors.
- Lead Time: Forecasts are generally more accurate for shorter lead times. A 6-hour forecast will typically be more skilful than a 72-hour forecast. Verification statistics are often presented as a function of lead time, showing the decay of skill over time. This helps users understand the inherent uncertainty as they look further into the future.
- Geographic Sector/Region: Forecast models perform differently over various terrains and regions. For instance, wind forecasts over flat, open terrain might be more accurate than those over complex mountainous regions or coastal areas with strong local effects. Verification is often conducted for specific regions (e.g., Ireland, the Irish Sea) to ensure local relevance.
- Parameter: Wind speed, gust, and direction each have their own verification challenges and skill levels. Gust forecasts, being inherently more turbulent and dependent on sub-grid scale processes, often exhibit lower skill than mean wind speed forecasts. Probabilistic forecasts for rare events also require specific verification approaches.
Operational users of The Wind Agent, such as crane operators or drone pilots, often require high accuracy for specific parameters (e.g., gust speed at height) and for short lead times (e.g., the next 6-12 hours). Verification programmes are tailored to these specific needs, providing targeted feedback on model performance.
This chart illustrates how different models can vary in their predictions, especially at longer lead times, highlighting the importance of ensemble verification.
07Why we publish track records
The Wind Agent is committed to transparency and providing users with the most reliable information possible. This commitment extends to publishing and making accessible the track records of the underlying forecast models. Publishing track records, based on rigorous verification, serves several critical purposes:
- Builds Trust: By openly sharing how well (or not so well) forecasts have performed, users can develop a realistic understanding of forecast limitations and build trust in the instrument's capabilities.
- Informs Decision-Making: Understanding the typical errors and biases of a forecast system helps users interpret the current forecast with appropriate caution. For instance, if a model consistently under-predicts strong winds in a particular scenario, users can factor that into their risk assessments.
- Highlights Strengths and Weaknesses: Verification data allows users to identify conditions or lead times where a particular model performs exceptionally well, or where it struggles. This knowledge empowers users to leverage the strengths and mitigate the weaknesses.
- Drives Improvement: For us, continuous verification is a cornerstone of our development programme. By systematically measuring forecast performance, we can identify areas for improvement, test new model configurations, and refine our post-processing algorithms to enhance accuracy and utility.
- Compliance and Accountability: In sectors where wind conditions are critical for safety and operations, verifiable forecast performance can contribute to compliance with internal safety standards and external regulatory requirements. The Wind Agent's 'Evidence Records' feature allows you to capture and store forecast data, providing an auditable trail for operational decisions.
Our commitment is to provide a precise instrument, not just a forecast. Verification is the scientific method behind that precision, ensuring that every number presented has a known and quantified level of reliability.
Questions
What is the difference between accuracy and skill in forecast verification?
Accuracy refers to how close a forecast is to the observed reality, often measured by metrics like MAE or RMSE. Skill, on the other hand, measures how much better a forecast is compared to a simple reference forecast, such as persistence (assuming current conditions continue) or climatology (assuming average conditions for the time of year). A forecast can be accurate but have low skill if a very simple forecast would have performed almost as well.
Why is forecast verification important for operational decisions?
Forecast verification is crucial for operational decisions because it provides an objective measure of forecast reliability. Knowing the typical errors, biases, and skill of a forecast helps operators understand the level of confidence they should place in the prediction. This enables better risk assessment, resource allocation, and adherence to safety protocols, especially when operating near wind speed or direction thresholds.
How does The Wind Agent use forecast verification?
The Wind Agent uses continuous forecast verification to monitor the performance of its underlying meteorological models against real-world observations. This data informs the development of our post-processing algorithms, helps us understand model strengths and weaknesses in different conditions, and underpins the reliability of features like the Exceedance Fan and Agreement Spine, ensuring the probabilities and comparisons presented are as accurate as possible.
Can verification tell me which model is 'best'?
Verification can identify which model performs 'best' for specific parameters, lead times, and regions, according to defined metrics. However, no single model is universally 'best' across all scenarios. One model might excel at short-range wind speed forecasts over land, while another might be superior for long-range gust predictions over the sea. The Wind Agent's approach is to provide an ensemble view and verification data, allowing users to assess performance contextually.
What is a 'proper scoring rule' and why does it matter for probabilistic forecasts?
A proper scoring rule is a method for evaluating probabilistic forecasts that incentivises forecasters to state their true beliefs or probabilities. The Brier Score is an example. If a scoring rule is not 'proper', a forecaster could improve their score by deliberately misrepresenting their forecast probabilities (e.g., making them more extreme than their true belief). Proper scoring rules ensure that the best strategy for a forecaster is always to be honest about their uncertainty, leading to more reliable probabilistic outputs for users.
SOURCES
- WMO Guide to Meteorological Instruments and Methods of Observation (WMO-No. 8)
- ECMWF Forecast Verification
- NOAA's Meteorological Verification Program
- Met Éireann: About Our Forecasts
- Stull, R.B. (1988). An Introduction to Boundary Layer Meteorology. Kluwer Academic Publishers.
Thresholds on this page are commonly cited figures, attributed to their source — never statutory limits. Modelled forecasts are planning support, not on-site measurement.