tailtwist · guided generation of rare weather
Guidance Shifts the Tail and Narrows It
A diffusion weather emulator can be steered towards hundred-year heatwaves and hurricane-force winds while it samples, and steering raises the hit rate: at the strongest setting the target event turns up in a steered draw about 90 times more often than its once-in-a-hundred base rate. tailtwist is a new tail-sampling strategy (preprint and GitHub repository to follow). measures what that steering does to the distribution the extra events come from, using an extreme-value engine the emulator has never seen.
Under guidance the upper tail moves to higher temperatures and narrows, more so at the stronger setting. The sampler concentrates the distribution on its target and does not reach further into the tail, so each steered draw's correction weight has to be checked before it is used.
the technical version
Both guided scales are exploratory (81 steered draws each, 23409 pooled values), so the curves stop at P99.9, and the narrowing in the headline is the P90 to P99.9 width read off them. The fitted shape is a separate quantity. A generalised Pareto fitted above the unguided pool's threshold gives ξ −0.265 unguided and −2.005 at the strongest scale, but by then 93% of the pooled values clear that threshold. That fit describes the bulk of a distribution the guidance has pushed past the cut, so its fall in ξ reflects that displacement and says nothing about the tail's length. The reliability diagnostics are below and on the application page.
×90
how much more often the target event turns up at the strongest guidance
×0.59
upper-tail width (P90 to P99.9) at the strongest guidance, relative to unguided (exploratory, n = 81)
The short version is tailspec's 02 · scale. This page has the method in full: the sampler configurations, the per-event fits, the reliability diagnostics, and what a reinsurance layer priced off these samples would have cost.
one · the problem
The expensive part of climate risk
A one-in-a-hundred-year event occurs, on average, once in a hundred simulated years. To study the tail you either wait, or generate enormous numbers of samples and discard nearly all of them.
Guided generation is the usual shortcut: nudge the model while it generates, and extreme events appear in nearly every draw instead of almost none. Each steered draw then has a correction weight (an odds ratio) so that averages over the steered sample still estimate the un-steered climate. In this project one guided draw cost the GPU time of roughly 153 plain ones.
two · the method, step by step
Guided Sampling in Six Steps
Turn the guidance up and the draws crowd past the event threshold, while the weights meant to keep the statistics unbiased degrade. Click through the steps and press each guidance strength on every one.
Run the emulator freely and the target event, a hundred-year heat over the target region, turns up in about one draw in a hundred. The tail is there, but reaching it is expensive.
Press a guidance strength to steer the draws.
1%
of free draws reach the event: the base rate that defines it
Each strength is its own fan-out at a different guidance scale. None of this is one sample set reweighted after the fact.
In the configuration measured here the rare events arrive and the correction weights fail.
three · the referee
How the comparison works, not what it found for any one event. The measured numbers checking the model's tail against ERA5 are on the pricing page.
Grading the tail from outside
tailtwist does not let the generative model grade its own tail. The verdict goes to tailspec, a separate extreme-value validation engine that fits the tail of the emulator's output and the tail of the ERA5 reanalysis record with the same classical statistics, and compares the two with no access to how either was produced.
four · the arithmetic
The equal-cost arithmetic
4909 < 5001
the guided probe's worth in unguided draws, against draws already on disk
1.8%
short of the free comparator, before counting the failed weights
The two sides are matched on GPU cost, since matching on draws would ignore the steered method's extra cost.
Converted into plain draws, the guided probe's entire GPU bill buys 4909 of them, fewer than the 5001 the archive already held, so on this probe the guided method did worse than not using it at all.
five · the other use
What the machinery is good for
The emulator underneath the steering is still a large-ensemble machine, and run plainly, with no guidance and no correction weights, it does what large ensembles are for: it answers counterfactual questions by drawing the same summer thousands of times under a different boundary condition.
Warming the sea-surface temperatures the model is conditioned on by a uniform 2 K and drawing 2500 fresh members makes the area-wide July heat event over Europe 1.75 times as likely as it is in the unperturbed pool of 2467: 3.5% of members clear the threshold against 2.0%.
×1.75
likelihood of the area-wide Europe heat event under a 2 K warmer sea surface (1.08 to 2.84, conservative)
1 of 4
scenario-and-threshold pairs whose interval clears 1: the mild storyline and the single-city threshold do not separate from noise
what this is not
The perturbation is a uniform offset applied to the whole prescribed sea-surface field, with no coupled ocean, no sea-ice response and no pattern scaling, and the estimand throughout is the emulator's own density rather than the real climate. The scenarios are drawn on independent seeds, so these are two binomial proportions compared, not a paired difference.
That is a response to the boundary condition alone, not a description of a 2 K warmer world. The generator damps the perturbation by roughly half, so a 2 K sea surface gives about 0.85 K of realised Europe-box warming. Compare it with the 0.83 K tail shift the emulator produces between its own early and late decades with no perturbation at all, which the storyline ladder brackets.
The design has 2500 members, one changed boundary condition, and a threshold fixed on the unperturbed pool and applied unchanged to both sides. It uses no steering and no odds ratios, so none of the reliability problems elsewhere on this page arise.
six · where this goes
The applied version
The same measurement applied to reinsurance layers and weather hedges is on the pricing page.
The emulator studied here is NVIDIA's cBottle, run through earth2studio, but the measurement applies to any diffusion emulator with a guidance term.
The code is private for now while the work is written up. I am always happy to talk through the method, the results, or the engineering: hello@wienkers.com, more at wienkers.com and tailspec.wienkers.com, where the evaluation engine behind this page lives and where 02 · scale is the short version of the finding here.