tailtwist · applied to a pricing desk
Only Unguided Draws Give a Usable Price
Two contracts are priced here off the emulator and off the ERA5 reanalysis: an excess-of-loss reinsurance layer on hurricane wind, and a cooling-degree-day hedge on a hot European summer. Guided draws reach the extreme events far more often, but their correction weights are degenerate, so every price on this page comes from unguided draws.
| unguided | guided | |
|---|---|---|
| draws reaching the heat event | 1% | 90% |
| correction weights | not needed | k = +∞ (heat), 0.96 to 5.82 (hurricane) |
| effective draws after weighting | all | 1.0 of 16 |
| price | −92% vs ERA5 on the 10/100-year wind layer at NE Caribbean, the widest gap of 3 sites | none |
one · the price gap
the technical version
The season-window check uses one illustrative site. A basin-wide claim needs the regional average in the chart, which pools the full year rather than only hurricane season and so complements the site-level numbers.
Pricing the Hurricane Layer
An excess-of-loss layer attaches at the 10-year wind level and exhausts at the 100-year level: the reinsurer pays for wind between the two. At every anchor the emulator's 100-year wind is lower than ERA5's (a regional-mean gap of −4.1 m/s), so the layer prices far cheaper than the observed record supports.
TC 10 m wind, 10/100-yr XoL layer, September leg. Whiskers = model 95% bootstrap CI for fit uncertainty only: it assumes independent samples, so it understates the true uncertainty given how correlated the underlying grid cells and time steps are.
NW Bahamas
(25.0N, 75.0W)
mispricing -59%
NE Caribbean
(18.0N, 65.0W)
mispricing -92%
N Gulf of Mexico
(28.0N, 88.0W)
mispricing -48%
The gap is sensitive to the season window. Restricted to August-October, the NW Bahamas anchor flips sign to +13%.
two · the weather desk
the technical version
The model's own draws are built from sea-surface conditions spanning 1979 to 2022, while the ERA5 comparison record covers only 2020 to 2023. Over that longer stretch the emulator's hot extremes have drifted warmer, by an amount that depends on how extreme the day is. Part of each city's gap may therefore come from the mismatch in time windows and not from a genuine difference in tail behaviour.
Cooling-Degree-Day Hedges in Four Cities
A cooling-degree-day (CDD) contract pays out one unit for every degree-day a summer runs above a strike. At the everyday strike the two sides roughly agree, and at the extreme strike the emulator pays a fraction of what ERA5 pays, because its hot tail is too short. Choose a city in the chart.
three · the evidence, cell by cell
The emulator's tail against ERA5
This is the gridded picture behind the prices, cell by cell for both perils. It opens on the bias map the pricing sections inherit, and switching variables shows each side's fitted tail on its own.
loading gridded return-level data…
why there is no guided quote
Degenerate Correction Weights
At the strongest heatwave scale a single draw has 99.99% of the total weight.
Integrating the weights more finely does not help, so the failure is in the estimator itself. Re-running the same 40 seeds at the 36 integration steps of the original paper, twice the 18 used here, moves k from 0.96 to 1.00, both far past the 0.38 up to which weights from 40 draws stay usable, and the effective sample size stays at 8% of the draws. The spread of the weights does not measurably change either (0.85 to 1.07 on a paired bootstrap, an interval containing 1).
how this was built
Built as two quarantined environments: a GPU sampling side that only writes zarr, and a CPU evaluation side that never imports the emulator. Evidence base: 20 000 unguided draws per peril plus the guided probes above. Dashboard numbers generated 2026-07-31, revision 3c92713. Reproduce the figures with python scripts/phase35_figures.py. The method itself is on the landing page.