01 · Setup
Four runs, one endpoint
Four HRRR cycles are tracked independently and all end at the same valid time, Day +2 12Z. Every hour drawn below is a real forecast — the 12Z run goes out to F48. The shaded band is the intersection of all four coverages, which is exactly the 06Z run in full: 10,582 valid times per cycle across 454 complete forecast groups, where the cycles can be compared on identical weather.
HRRR and MRMS go through the same PyFLEXTRKR configuration as two separate streams, so neither side is privileged. Only the resulting track objects meet, at the matching step. The domain is 25–51°N, 110–70°W. Of the 1,836 initializations evaluated, 19 contain no MCS anywhere in the domain at any lead.
02 · Method
How we pair up objects
- cd is the centroid separation, md the minimum distance between object boundaries. Neither term is clamped at zero, so a large centroid separation can still veto a pair whose edges overlap.
- Candidate match when TI > 0.2, greedy one-to-one assignment.
- Matched pair = hit · unmatched forecast = false alarm · unmatched observation = miss.
- Identical hourly valid times, so the temporal term drops out.
- Baseline scales cd max = 250 km and md max = 100 km — far larger than Skinner's storm-scale 40 km, because MCSs are. These are settings, not constants: section 03 re-runs everything with cd max from 100 to 250 km.
Across all 73,842 matched pairs the median centroid separation is 157 km and the median boundary separation is 0.3 km — so a typical matched pair has centroids far apart but edges almost touching, which is what the boundary term is there to catch. Method follows Skinner et al. (2018, Weather and Forecasting). There is no object analogue of correct negatives, so CSI is punished from both sides. Note this is not nearest-centroid matching: all forecast × observed pairs are scored, candidates above the threshold are sorted by total interest, and assignment is greedy and one-to-one — a pair can win on boundary overlap even when its centroids are not the closest.
Why CSI sits near 0.31 when a case looks fine
HRRR produces two to three MCS objects almost regardless of how many are really there. That single fact drives most of the score.
| Observed MCSs | Hours | Forecast (mean) | CSI | Bias |
|---|---|---|---|---|
| none | 6,851 (10%) | 1.93 | — | ∞ |
| 1 | 20,828 (30%) | 2.15 | 0.259 | 2.15 |
| 2 | 23,114 (33%) | 2.66 | 0.335 | 1.33 |
| 3–4 | 17,130 (25%) | 3.27 | 0.357 | 1.00 |
| 5+ | 1,778 (3%) | 3.54 | 0.335 | 0.68 |
scene_complexity_scores.csv · 69,701 valid forecast hours.
- Sparse hours are over-forecast: with one observed object the bias is 2.15, and in 10% of hours nothing is observed at all while HRRR still makes about two.
- Busy hours are under-forecast: with five or more observed, the bias flips to 0.68 and objects are missed instead.
- The overall bias of 1.38 is the average of two opposite errors, not a mild uniform over-forecast.
- Multi-object scenes score better (CSI 0.34–0.36) than single-object ones (0.26). Case studies come from active hours, so they look better than the seasonal mean.
03 · Matching scale
How much of the score is the 250 km tolerance?
More than everything else in this report put together. Re-running all three seasons with a tighter centroid scale, every other setting held fixed, moves CSI further than ageing a forecast from F01 all the way to F48 does.
| cd max | POD | FAR | CSI | Median pair distance |
|---|---|---|---|---|
| 100 km | 0.284 | 0.794 | 0.136 | 92 km |
| 150 km | 0.416 | 0.698 | 0.212 | 122 km |
| 200 km | 0.502 | 0.635 | 0.268 | 143 km |
| 250 km | 0.559 | 0.594 | 0.308 | 157 km |
data/cdmax_sensitivity/. Frequency bias is 1.376 in all four runs —
the object counts never change, only which pairs are allowed to match.
CSI more than doubles from 0.136 to 0.308 as the matching scale goes from 100 to 250 km.
- cd max is not just a cut-off, it is the normalisation of the centroid term. Halving it both tightens which pairs survive and steepens the penalty on the ones that do, which is why the effect is so large.
- The absolute CSI is therefore not a model property on its own. It is a joint statement about the model and the tolerance we allow.
- Comparisons within this bundle — between cycles, leads, years or life-cycle classes — are safe, because they all share one scale.
- Comparisons against any published CSI are not, unless that study used the same centroid and boundary scales.
04 · Skill vs forecast lead
Slow decay, thinning sample
CSI falls from 0.35 to 0.26 across F01–F48. Frequency bias starts very near one (0.98 in the first six hours), jumps to a peak of 1.55 around F19–F24 and then eases back. Position error grows steadily from 111 km to 196 km, and size bias climbs from 1.17× to 1.91×.
Object skill
Frequency, size and position bias
Hover any point for the sample behind it. Dashed past F30, where fewer than four cycles contribute. Position error uses the right-hand axis in kilometres; the other two are forecast/observed ratios against the left axis.
The first six hours are the interesting part. At F01–F06 the model is close to unbiased in object count and only 111 km off in position, and CSI is at its highest. Everything after that is the model drifting away from its initial state — by F19–F24 it is making about 1.55 objects for every observed one, and the typical matched pair is 163 km apart.
Read the long-lead end carefully. These bins pool all cycles at a given lead, not at the same valid time. Past F30 the sample collapses from four cycles to three, two, then one, and the verifying time of day shifts with it. So the decline mixes forecast ageing with sampling and the diurnal cycle — it is not a clean measure of any one of them. Section 06 holds the weather fixed instead.
05 · Life cycle
Two opposite errors hiding inside one curve
Split the objects by whether the MCS already existed when the forecast started, and the flat-looking lead-time curve comes apart into two trends running in opposite directions. This is the clearest new result in the three-season evaluation.
CSI by life-cycle class
Frequency bias by life-cycle class
Matching is performed separately within each class, so these are two independent evaluations rather than a partition of one score. Observed objects are classed pre-existing when their first robust-MCS time is at or before initialization; forecast objects use first robust-MCS time at or before F01 as a proxy.
| Class | Objects | POD | FAR | CSI | Bias |
|---|---|---|---|---|---|
| Newly initiated | 98,336 obs | 0.459 | 0.665 | 0.240 | 1.37 |
| Already ongoing | 33,756 obs | 0.498 | 0.643 | 0.262 | 1.40 |
mcs_objects_by_lifecycle_summary.csv. Pooled over all leads the two
classes look almost identical — which is exactly why the split by lead matters.
Ongoing MCSs: CSI 0.36 → 0.12, bias 0.67 → 3.30.
New MCSs: CSI 0.12 → 0.26, bias flat near 1.4.
- HRRR is good at systems it inherits and bad at keeping them honest. At F01 it tracks ongoing MCSs well (CSI 0.36) but has too few of them (bias 0.67).
- By F30 the observed pre-existing population has fallen from 2,709 objects to 230 — a 92% decay — while HRRR's has only fallen from 1,824 to 760, a 58% decay. The model will not let long-lived MCSs die.
- Newly initiated systems run the other way: terrible at short lead (CSI 0.12 at F01, where there is almost nothing to initiate yet) and steadily improving as the forecast has time to spin up convection.
- Because the two curves cross near F08–F10, the pooled score looks much flatter than either component. Any downstream hazard model inherits both errors, not the average.
Two cautions. The forecast-side class is a proxy — an MCS that is robust by F01 is assumed to have been inherited rather than generated, which cannot be verified from the forecast alone. And past F42 the pre-existing sample is down to a few dozen objects, so the far-right end of the bias curve is noisy even though its direction is already established well before then.
06 · Initialization cycle
Does the cycle matter?
On the matched-valid-time sample every cycle sees identical weather — the same 10,582 valid times — so only forecast age differs. With three seasons behind them the bootstrap intervals are now narrow enough to separate the cycles, which they were not on one month of data.
| Cycle | Lead range | POD | FAR | CSI | 95% CI |
|---|---|---|---|---|---|
| 12Z | F19–F48 | 0.561 | 0.636 | 0.283 | 0.271–0.295 |
| 18Z | F13–F42 | 0.568 | 0.614 | 0.298 | 0.285–0.311 |
| 00Z | F07–F36 | 0.599 | 0.588 | 0.323 | 0.309–0.336 |
| 06Z | F01–F30 | 0.538 | 0.561 | 0.319 | 0.305–0.333 |
10,582 identical valid times per cycle · 454 complete forecast groups. Intervals resample whole forecast groups, not single hours.
The ordering is 00Z ≈ 06Z > 18Z > 12Z, spanning ΔCSI = 0.040 from best to worst.
- 12Z is clearly worst. Its interval [0.271, 0.295] lies entirely below those of 00Z and 06Z. It is also the oldest forecast in this comparison (F19–F48), so this is largely forecast age rather than anything special about 12Z.
- 00Z and 06Z are indistinguishable — their intervals overlap almost completely — even though 06Z is a full six hours younger. Freshness stops buying skill somewhere around F07.
- 18Z overlaps 06Z only marginally, so that pair is suggestive rather than settled.
- Note what POD and FAR do here: the newest cycle (06Z) has the lowest POD but also by far the lowest FAR and the lowest frequency bias (1.23 against 1.54 for 12Z). Older forecasts do not miss more — they over-produce more.
07 · Position error
Shifted north, stretched southwest–northeast
Each point is one matched pair: the forecast centroid minus the observed centroid. Two questions, and on three seasons both now have an answer. Is there a net shift? Yes, a modest one — forecasts sit about 24 km too far north. Is the scatter round? No: the cloud is stretched along a southwest–northeast axis.
| Cycle | Pairs | Median north | Major axis | Axis ratio |
|---|---|---|---|---|
| 12Z | 22,758 | +25.4 km | 32.6° | 1.20 |
| 18Z | 19,713 | +20.8 km | 36.3° | 1.19 |
| 00Z | 17,868 | +21.2 km | 37.3° | 1.21 |
| 06Z | 13,503 | +28.3 km | 37.8° | 1.21 |
| all | 73,842 | +23.6 km | 35.7° | 1.20 |
mcs_matched_centroid_offsets_summary.csv. Angles are counter-clockwise
from east, so 45° would be exactly southwest–northeast.
A northward displacement of about 24 km, on a cloud of error stretched along 36° with 151 km of spread along the axis against 126 km across it.
- The northward shift is small next to the 157 km typical displacement, but it is consistent: every cycle is north, and 58% of all pairs are north of their observed counterpart.
- The east–west offset is essentially zero (median −3 km, 49% east).
- The anisotropy is systematic, not sampling: all four cycles give 33–38° with the same 1.20 ratio.
- That axis is roughly the warm-season steering direction, which points to along-track error — the system run too fast or too slow down a broadly correct path rather than displaced to one side of it. That last step is interpretation; confirming it needs track motion vectors, which these pair records do not carry.
What changed from the one-month evaluation. On June 2021 alone the net shift looked like zero and the axis ratio looked stronger (1.31). With 73,842 pairs instead of 5,949 the ratio settles at 1.20 and a real northward bias emerges from the noise. The direction of the tilt is the part that held up.
08 · Structural error
Too small, too intense
This is the strongest signal in the evaluation. Every value below is a forecast-to-observed ratio, so 1.0 would be perfect.
| Cycle | Cloud shield | Rain area | Rain rate | Volumetric |
|---|---|---|---|---|
| 12Z | 2.19× | 0.67× | 1.53× | 1.44× |
| 18Z | 2.08× | 0.64× | 1.55× | 1.36× |
| 00Z | 2.08× | 0.66× | 1.58× | 1.37× |
| 06Z | 1.91× | 0.61× | 1.56× | 1.28× |
Median forecast/observed ratios ·
mcs_matched_property_violin_summary.csv.
- Every cell has the same sign in every cycle. Nothing here depends on which initialization you look at.
- The cloud-shield error grows with forecast age — 1.91× at 06Z (youngest) up to 2.19× at 12Z (oldest) — while the rain-rate error does not move at all.
- Rain rate is the most stable bias in the whole evaluation: 1.53–1.58× across every cycle.
HRRR is not simply too wet or too dry. It concentrates too much intensity into too small a rain area, underneath a cloud shield that is too broad. The rain-area and rain-rate errors have opposite signs and partly cancel in the volumetric total — which is exactly the kind of predictor shift that would move a downstream hazard model even when total rainfall looks reasonable.
09 · Year-to-year stability
Does any of this move between seasons?
Three warm seasons let us ask whether these are model properties or properties of one summer. The skill scores barely move; the frequency bias does.
| Year | Hours | POD | FAR | CSI | Bias |
|---|---|---|---|---|---|
| 2021 | 23,012 | 0.562 | 0.572 | 0.321 | 1.31 |
| 2022 | 23,436 | 0.527 | 0.599 | 0.295 | 1.31 |
| 2023 | 23,253 | 0.591 | 0.609 | 0.308 | 1.51 |
mcs_objects_by_year.csv.
CSI spans only 0.295–0.321 across three seasons, but frequency bias jumps to 1.51 in 2023.
- CSI is stable to about ±0.013 around 0.308 — smaller than the spread between initialization cycles, and far smaller than the matching-scale effect in section 03.
- 2023 is the odd season: the highest POD and the highest FAR. HRRR simply made more objects that year — 62,841 against 58,509 in 2021 — while observed counts fell, from 44,506 to 41,571.
- So the over-forecasting tendency is not fixed. It is the part of this evaluation most likely to shift with model version or season.
10 · Case study · 20 June 2021
The same bias in one storm
One of the 1,836 initialization sets, shown in full. All four forecasts put an MCS in roughly the right place at 12Z, so the spatial envelope is recognizable. But every forecast mask extends well beyond the observed object, with area ratios from 1.75× to 2.24× — the aggregate cloud-shield bias, in a single case.
Take-home messages
What we know so far
- Overall MCS properties. The structural error is the strongest signal and the one most likely to matter downstream. Cloud shields are 2.08× too large, rain areas 0.63× too small, and mean rain rates 1.54× too high. The last two have opposite signs and partly cancel, leaving volumetric rain at 1.37× — so a model can look acceptable on total rainfall while its rain area and intensity are both wrong. Every cycle, every year, same sign.
- Life-cycle comparison. The single most useful split in the three-season data. HRRR tracks inherited MCSs well at short lead (CSI 0.36 at F01) but will not let them dissipate — observed pre-existing objects fall 92% by F30 while forecast ones fall 58%, driving frequency bias from 0.67 to 3.30. Newly initiated systems run the opposite way, from CSI 0.12 at F01 up to 0.26. The pooled curve is flat only because these two cancel.
- Lead-time comparison. Skill decays slowly: CSI 0.35 → 0.26 over F01–F48, position error 111 → 196 km, size bias 1.17× → 1.91×. Frequency bias is near one for the first six hours, peaks at 1.55 around F19–F24, then eases. Past F30 the sample collapses from four cycles to one, so the long-lead decline mixes forecast ageing with sampling and the diurnal cycle.
- Initialization comparison. On identical weather — the same 10,582 matched valid times — the ordering is 00Z ≈ 06Z > 18Z > 12Z, spanning ΔCSI = 0.040. With three seasons the bootstrap now separates 12Z from both 00Z and 06Z, which one month could not. But 00Z and 06Z remain indistinguishable despite six hours of age difference, so freshness stops paying off somewhere around F07. The newest cycle wins on FAR and bias, not on POD.
- How much to trust the absolute number. CSI runs 0.136 to 0.308 depending only on the matching scale, so treat 0.308 as a statement about this configuration, not a model constant. Year to year it is stable to ±0.013; the frequency bias is not, rising to 1.51 in 2023.
Open question for discussion: how do we handle these biases before trusting the hazard output — and should the life-cycle split be carried through into the hazard model itself, given how differently the two classes behave?
April–August 2021–2023 · 1,836 initializations · 69,701 valid
forecast hours · 73,842 matched object pairs.
Overall POD 0.559 · FAR 0.594 · CSI 0.308 · frequency bias 1.376.
Matching: cd = 250 km, md = 100 km, TI > 0.2, Skinner et al. (2018).
Figures and data in assets/, data/ and scripts/.
Presentation version →