tailtwist · guided-generative rare-event sampling, audited from outside
Guidance buys samples. The odds ratio buys unbiasedness. Neither buys a tail you can price.
NVIDIA's cBottle diffusion emulator draws a global atmosphere in seconds, and its guided-generative extension promises what rare-event work actually needs: draws steered into the tail, carrying an odds ratio that reweights them back onto the model's own density. We audited that promise from the outside, using tailspec as the extreme-value engine, and then asked the question a sampler cannot answer about itself. What would you quote off this tail, and would you pay for the guided version of it?
−92%
how far the emulator's own price for a 10/100-year hurricane-wind layer at NE Caribbean falls below the same layer priced off ERA5
k̂ = +∞
the odds-ratio importance weights at both probed heatwave guidance scales, with the guidance functional targeting the scored event itself
4909 < 5001
unguided draws bought by the whole guided probe at equal GPU cost, against the unguided draws already sitting in the store (1.8% short)
Read every number here as a statement about cBottle's own density, never about reality. That is the estimand a self-contained sampler diagnostic is stuck with, and it is exactly why the ERA5 comparison below has to come from an engine the emulator had no hand in.
one · the price
What the emulator would quote, and what ERA5 says the same layer costs
Start where a sampler's tail actually gets spent. We priced a 10/100-year excess-of-loss layer on 10 m wind at the three pre-registered North Atlantic anchors, fitting a generalised Pareto tail to each side and integrating the same fixed layer over each fitted survival function. Both sides are restricted to September and to the same 120 observations per year, so the emulator is compared against ERA5 on identical terms and never across months. Premiums are per-observation expected losses, not annualised, which is why they read as small numbers: the mispricing fraction, not the level, is the decision-relevant quantity.
| Anchor | ERA5 | cBottle (unguided) | Model 95% CI | Mispricing | ASO leg |
|---|---|---|---|---|---|
| NE Caribbean | 2.83e-3 | 2.38e-4 | 2.00e-4 to 2.82e-4 | −91.6% | −89.9% |
| NW Bahamas | 2.12e-3 | 8.60e-4 | 7.54e-4 to 9.84e-4 | −59.4% | +13.0% |
| N Gulf of Mexico | 1.42e-3 | 7.36e-4 | 6.63e-4 to 8.14e-4 | −48.1% | −99.7% |
The emulator under-prices the layer at all three anchors, from −48% to −92%. Phase 1 taught us not to generalise a single pooled anchor to an area claim, so the area-level statement comes from the gridded region mean instead, and it agrees: a 100-year return level of 15.42 m/s against ERA5's 19.53 m/s, a gap of −4.12 m/s over 5844 ERA5 steps. That region mean pools the whole year at the Phase-1 rate rather than the September layer's, so it is the independent area-level check and not a recomputation of the table above, and its levels are not comparable with the anchors' attachment points. A short tail prices a cheap layer, and the reinsurance arithmetic converts the shortfall into money without any further modelling assumption.
Phase-1 Gate-1: never generalise a single pooled anchor to an area claim -- report all three anchors (nw_bahamas, ne_caribbean, n_gulf_of_mexico) and quote region_mean_return_levels (the gridded region mean, all-year rate, from phase1_tail_audit.json) for any area-level statement. region_mean_return_levels is copied verbatim as context; it is NOT September-conditioned like the anchors above (it pools the full year at the Phase-1 1460 obs/yr rate), so it is not directly comparable to the anchor blocks' return levels -- read it as the independent area-level check, not a recomputation at the same rate.
TC 10 m wind, 10/100-yr XoL layer, September leg. Whiskers = model 95% L-moments bootstrap CI (fit uncertainty only, iid assumption -- see the CI caveat in pricing_tc.json's notes).
NW Bahamas
(25.0N, 75.0W)
mispricing -59%
NE Caribbean
(18.0N, 65.0W)
mispricing -92%
N Gulf of Mexico
(28.0N, 88.0W)
mispricing -48%
The direction is robust, the magnitude is not. Repeating the identical construction on August to October at 368 observations per year moves every anchor, and at NW Bahamas it flips the sign outright (−59% in September against +13% across the wider season). Any single anchor, on any single seasonal window, would have supported a confident and wrong headline. Three anchors and two windows support a narrower one, which is the claim we make.
cBottle's own unguided density, not reality. Every premium below is the price you would quote off each tail *fit* (model or ERA5), not a real-world premium; the external ERA5 comparison is what turns the model-relative gap into a decision-relevant mispricing number.
The same question on a weather desk
Cooling-degree-day contracts price the other peril in this project. For each of four European cities we took the ERA5-implied p90 and p99 CDD strikes (base 18°C) and paid them off each side's own climatological distribution, 5001 model draws against 1472 JJA ERA5 steps, the summer subset of the four-year 6-hourly record the note below quotes in full. Every city comes in short at the p99 strike, from 67.7% below the ERA5 payout at Amsterdam Schiphol to 99.1% below it at London Heathrow. It is the same missing tail as the hurricane layer, expressed as a weather derivative that settles nearly worthless.
The unguided fan-out spans SST years 1979-2022 while the ERA5 t2m timeseries pools only 2020-2023 (5844 6-hourly steps, a 4-year record) -- read every model-vs-ERA5 gap below alongside the emulator's own measured late-minus-early JJA tail shift, +0.81 K (p90) / +0.83 K (p99) (Phase-1 measurement, Europe region). The p90 strike is exactly the level the payout at that strike integrates, so the +0.81 K figure is the more directly relevant of the two for reading the p90 mispricing; the +0.83 K figure is the relevant one for p99.
cBottle's own unguided JJA density, not reality. Every payout below is what you would pay off each side's own CDD *distribution* (model draws or ERA5 steps), not a real-world settlement value; the external ERA5 comparison is what turns the model-relative gap into a decision-relevant mispricing number.
The guided method does not get to quote
Every price above comes off unguided draws. That is not a modelling preference. The guided-plus-odds-ratio estimator trips the trust gate on the heatwave probe the gate was pre-registered for, and the same two-clause rule, applied post-hoc to the tropical-cyclone ladder, trips there too, so no guided estimate in this project carries a standard error we would propagate into a price, and a premium without an uncertainty is not a quote. The audit below is the account of how a method that enriches the tail by a factor of tens still ends up with nothing a desk can use.
No estimator standard error exists to propagate into price uncertainty. The pre-registered trust gate (gate.rule, generalises_gate2) was computed by Phase 3 for the heatwave probe and fired True there (k-hat = +inf at both probed scales, khat_is_inf_both_scales=True). Applied POST-HOC to the Phase-2 TC ladder (gate.applied_to_tc -- the SAME two-clause rule, not Phase-2's own separate single-clause khat_passes_gate criterion), it also fires True: every TC scale independently clears the true-branch's khat_weights > max(good_k(n), 0.7) clause (raw-weight k-hat 0.96-5.82 across all four probed scales). Phase-2's OWN separate, single-clause gate (khat > good_k(n), no dominance term, threshold 0.376 at n=40) independently also fails at every TC scale -- corroborating context, not the criterion generalises_gate2 above evaluates. No guided+odds-ratio importance-sampling estimate in this project therefore carries a trustworthy SE, on either criterion. Every price quoted in results/phase35_pricing_tc.json and results/phase35_pricing_heat.json is computed off the UNGUIDED draws only -- never off a guided/odds-ratio-reweighted estimate; both files' own 'estimand' strings say so (pricing_leg_sources below, copied verbatim, both containing the word 'unguided').
two · the audit
Enrichment works. The weights do not.
A price is only as good as the tail underneath it, so before asking what guidance buys, look at what the unguided emulator already gives. The Phase-1 maps below hold the gridded extreme-value fit for both perils, model against ERA5, on a climatological axis (cBottle is free-running and boundary-conditioned, so nothing here is date-matched). This is the reference the layer prices were quoted off, and the surface the guided sampler was supposed to improve on.
loading gridded return-level data…
Guidance enriches, and it is not close
Steering works as advertised. At a softplus guidance strength of β = 5 K on the Paris-box areal-mean temperature, the fraction of draws landing above the pre-registered p99 event threshold rises from the unguided base rate of 1% to 45.7% at guidance scale 1200 and 90.1% at scale 3700, an enrichment of 45.7× and 90.1× on 81 draws per scale, with the box mean lifted by 6.11 K and 19.49 K respectively. As a proposal mechanism for rare events, that is exactly the behaviour the method promises: the sampler spends its compute where the event lives instead of waiting for one draw in a hundred.
This ladder is the FORWARD-ONLY production audit (n=81 draws/scale at scale_1200 and scale_3700, guided sampler pass only, no backward phases, no odds ratio -- each rung's own 'label' field records its exploratory n verbatim), NOT the full-odds-ratio probe in heatwave_measurements above (n=16/scale, forward + backward + backward-no-guidance passes). Per phase3_physics_audit_heat.json's own or_probe_note, these are two different sampling paths even before weighting is considered -- do not read the two tables as one ensemble at two different n.
The weights that were supposed to pay for it
Enrichment on its own biases everything. The odds ratio is what buys the bias back, reweighting each guided draw onto the model's unguided density, and the weights are therefore the load-bearing component of the whole method. We diagnosed them with PSIS, whose shape parameter k̂ states directly how many moments of the weight distribution exist. On the tropical-cyclone ladder, k̂ climbs from 0.96 at scale 128 to 5.82 at scale 1280, against a gate of 0.38 at n = 40. Every probed scale fails, and the effective sample size tells the same story in plainer units: 1.00 effective draws out of 40 at the strongest guidance, with a single draw carrying 97% of the total weight. Phase 2's own single-clause criterion and the two-clause gate below, applied to this ladder post-hoc, agree on it.
The obvious objection is that tropical-cyclone guidance only correlated with the scored event, so we built the closest-to-favourable case we could and ran it as a pre-registered generalisation test. In the heatwave probe the guidance functional targets the areal-mean temperature that the event threshold is defined on, so the guidance target is the estimand. The weights degenerate anyway. k̂ is infinite at both probed scales, and at scale 3700 one draw of 16 carries 99.99% of the weight, leaving 1.00 effective draws. The pre-registered gate fires. Degeneracy is not an artefact of a mismatched guidance target, however, and that is the finding: it survives the case built to be kindest to the method.
n=16 per scale bounds what can be said; 2 of 2 pre-registered scales had data and voted. verdict strength differs by scale: scale_3700 votes true on two independent diagnostics (khat_weights=+inf AND dominance=0.9999>0.5); scale_1200 votes true on khat_weights=+inf alone (its own dominance=0.277<0.5 and ESS/n=0.393<0.5 do not independently trigger the true branch). khat_weights=+inf itself rests on a 3-point GPD tail fit (arviz's psislw fits M=min(n/5, 3*sqrt(n))=3 tail draws at n=16) -- a real, if extreme, PSIS diagnostic, not missing data, but drawn from very few points. If khat_weights were instead treated as unavailable at scale_1200, that scale would satisfy neither the true- nor the false-branch (dominance<0.5 but ESS/n<0.5), making the overall verdict 'insufficient' rather than True. The pre-registered rule is applied verbatim to the actual computed values (True) -- this note records evidence strength, not a rule change.
What the guidance does to the tail it enriched
The weights fail loudly. The physics fails quietly, which is worse. Fitting the same generalised Pareto tail to the box pool at a single fixed threshold, the shape parameter ξ falls from −0.265 unguided to −0.661 at scale 1200 and −2.005 at scale 3700, while the exceedance rate at that same threshold rises from 6.6% to 52.1% and 93.5%. A strictly shorter, more sharply bounded tail is not what a physical intensification of heat extremes looks like. At the top rung almost the whole pool sits above the threshold, so that fit is no longer a tail fit at all, and the guidance pressure that bought the enrichment has swamped the distribution it was drawing from. Read the collapse as an artefact of extreme guidance, not as a physical result, and note the sample each guided rung rests on (exploratory: n=81).
As guided draws move from unguided through scale_1200 to scale_3700, the box-pooled GPD shape parameter falls steeply more negative (-0.265 -> -0.661 -> -2.005) while the SAME shared threshold's exceedance rate rises from 0.066 to 0.521 to 0.935: at scale_3700, 93% of the pooled values exceed the threshold, so that rung's 'fit' is not a tail fit at all -- it is fitting the GPD to the bulk of the distribution. A real physical intensification of heat extremes would not produce a shorter, more strictly bounded tail; read this ladder as an artefact of extreme guidance pressure swamping the pooled box (see or_probe_note above and scale_3700's own reading_note in phase3_physics_audit_heat.json), not a physical finding.
And then the arithmetic
Suppose you tolerated all of the above. The last comparison is the one that ends it. A full odds-ratio heatwave draw costs 1100 s of A100 time, measured on the probe job itself, two orders of magnitude more than an unguided draw at the Phase-1 measured base-model rate quoted below, so the honest comparison is at equal GPU cost rather than at equal sample size. The whole probe, 32 guided draws across both scales, buys 4909 unguided-equivalent draws. The unguided store already holds 5001. Paying for guidance here purchases 1.8% fewer draws than were free, and every one of the purchased draws would arrive with the degenerate weights above. The margin is thin, and it should be quoted thinly: the per-scale figure of 2454 paired with the 32-draw total would wrongly suggest a factor of two.
a full-OR heatwave draw costs 1100 s/sample -- MEASURED, job 26548566 (sacct: 32/32 full-OR samples, 0 rejects, 9.8 GPU-h total = 1102.5 s/sample, rounded to 1100 and applied flat across both scales, since the sacct total does not break the cost out per scale). This value is the seconds_guided fed to equal_cost_n for every scale below, replacing the earlier provisional forward-only x6 multiplier. For context only (not used in any cost computation here): WP-D's forward-only production fan-out measured 115.6 s/sample (job 26548553, 162/162 samples, both scales pooled, 5.2 GPU-h total) -- so the backward pass + backward-no-guidance phase + 3 Hutchinson probes atop one forward pass cost the OR probe roughly a further 9.5x on top of a forward-only draw, not the provisional ~6x guessed pre-measurement. Every event's monte_carlo_equal_cost block follows Phase 2's mc_estimate n_projected convention (tailtwist.weights.estimators.equal_cost_n, seconds_unguided=7.17 default, the Phase-1 WP1 measured base-model cost, left at its default here -- never overridden) -- the project's costing rule: compare standard errors at equal GPU cost, never at equal N.
None of this indicts the idea of guided rare-event sampling. It indicts this configuration of it, measured on this emulator, at these guidance strengths, with these sample sizes, and it does so against a rule written down before the numbers existed. The full gate, both scale tables and the refusal in its own words are on the rigour page.
Every number below describes cBottle's own unguided or guided+odds-ratio density, not reality (see each section's own 'estimand' string, copied verbatim from its source file); the external ERA5 comparison lives in the pricing legs (results/phase35_pricing_tc.json, results/phase35_pricing_heat.json), not here.
three · boundary-condition sensitivity
A dose-response ladder on the SST condition, not an attribution claim
What follows perturbs the emulator's SST conditioning field by a uniform +1 K and +2 K, and measures how cBottle's own July density over the Paris box responds. The generator damps that condition roughly 2:1 on its way into realised temperature (the +2 K condition realises +0.85 K of Europe-box temperature in the fan-out mean), so a rung is the response to a boundary-condition perturbation and never a "+2 K world". This is not event attribution. No real heatwave is being explained, and nothing here is a claim about reality's own sensitivity to sea-surface temperature.
The generator damps the SST boundary-condition perturbation ~2:1 (Task 6's paired canary: a +2 K SST condition realised ~+0.90 K of Europe-box t2m, ~0.45 K per K of condition). Read delta_1/delta_2 as the response to a +1 K/+2 K perturbation of the SST BOUNDARY CONDITION, never as a '+1 K/+2 K world' -- each scenario block above states its own measured realised_warm_shift_vs_delta0_K.
The ladder is worth running because it exercises the one thing an emulator of this kind is supposed to be good at, namely responding coherently to the boundary conditions it was conditioned on, and because it can be checked against a signal the model produces without any perturbation at all. The realised box-mean shifts are +0.524 K at +1 K (standard error 0.130 K) and +0.845 K at +2 K (0.130 K), against the +0.834 K late-minus-early contrast the emulator already carries on its own 1979-1990 to 2011-2022 trend axis. The ladder brackets that free contrast, but only just: the top rung clears it by 0.011 K, well inside its own sampling error, so read this as a plausibility check on magnitude rather than as comfortable straddling.
| SST condition | Event | Ratio | Interval | Excludes 1 |
|---|---|---|---|---|
| +1 K SST condition | box area p99 | 1.29 | 0.77 to 2.16 | no |
| +1 K SST condition | Paris point p99 | 0.99 | 0.56 to 1.75 | no |
| +2 K SST condition | box area p99 | 1.75 | 1.08 to 2.84 | yes |
| +2 K SST condition | Paris point p99 | 1.24 | 0.73 to 2.13 | no |
Of the four (scenario, threshold) likelihood-ratio conservative CIs, 1 of 4 exclude 1 and 3 of 4 have a point ratio >= 1. Read the ratios above 1 as a consistent DIRECTIONAL finding across both event definitions and both scenarios (larger delta always raises or holds the ratio), not as four independently statistically-significant results at this n -- only the largest-signal cell (area_p99 at delta_2) clears its own conservative interval at 95%.
One cell needs disclosing rather than defending. At the +1 K rung the Paris point event fires 46 times in 2467 factual draws and 46 times in 2500 perturbed draws, an identical count on both sides, so its ratio of 0.9868 is the ratio of the two pool sizes and carries no signal at all. That is what a single-grid-cell event at a base rate of 1.9% looks like at this sample size. The corroboration for the point event has to come from the quantile shift at the same cell, not from the crossing count, and the box-level event is the one carrying the directional finding.
The gate that mattered passed. Restricting Phase 1's unguided July draws and the dedicated zero-perturbation top-up to the same event definition gives consistent exceedance rates (combined estimate 0.01986 vs phase1-only Wilson CI [0.012634708986236521, 0.02557445711008145]), which is the only fatal check in this arm: without it the factual pool would not be a factual pool. The remaining checks are plausibility gates, and one of them reads FAIL for exactly the count-tie reason above.
- AMIP-style SST-only conditioning: no coupled ocean or sea-ice response to the perturbation, and no sea-ice adjustment channel exists in the conditioning input.
- The coarse ~100 km HEALPix-6 generator only; no super-resolution stage.
- Uniform-offset delta storyline: +delta_k K applied over the whole AMIP mid-month SST field, no pattern scaling. The AMIP tosbcs field is globally complete (zero NaN) -- it is NOT 'land-filled' in the sense of a separate fill step patching a land mask onto an ocean-only field; no such step exists, the source field itself carries SST-like values everywhere -- so encode_sst's land mask is inert and continental cells of the conditioning input are warmed too (no ocean mask). This mirrors the wording each archived sample carries in its own guidance.limits attrs (Task 6).
- cBottle's own unguided density throughout, never reality; climatological axis only (July mid-month SSTs across 1979-2022), never date-matched.
- Unpaired, independent seeds across scenarios (attribution_grid.py, by design): the likelihood ratios above compare two independent binomial proportions, not a paired difference.
- Year coverage is near-uniform per scenario (57 cycles of the 44-year 1979-2022 span, +/- 1 sample) -- a note, not a correction.
- TC-side attribution is out of scope: the SST response at this generator's resolution is not defensible for tropical-cyclone physics (heat only).
cBottle's own unguided density response to a perturbation of the SST BOUNDARY CONDITION, not real-world event attribution and not a claim about reality's SST sensitivity. The external ERA5 comparison (via the Task-3 CDD strikes reused below) is what turns the model-relative shift into a decision-relevant number.
four · how this was built
Methods, provenance and what it cost
Two environments, one contract
Sampling and evaluation never share a process. Generation runs on the GPU side (python 3.13, earth2studio 0.15.0 (cbottle extra), single A100 80 GB) against checkpoints pinned at revision eebd93c, and writes zarr. Evaluation runs on the CPU side (python 3.10, tailspec, no earth2studio import anywhere on this side) and reads it. The boundary is deliberately narrow, zarr plus JSON and nothing else, so that no evaluation result can quietly inherit a sampling-side assumption through a shared import. Every sample store must satisfy a frozen provenance contract before the evaluation side will look at it: seed, guidance configuration, the odds-ratio terms, sampler steps, Hutchinson probe count, package and checkpoint revisions. The contract only ever gains keys, never loses them, which is what makes an archived store from an early phase still readable by the last one.
Draws and seeds
The unguided fan-out is 20 000 draws per peril, from which the pricing legs use the conditioned subsets that match each contract (5001 JJA draws for the cooling-degree-day strikes, 481 474 pooled September values for the hurricane layer). The attribution arm draws its own grid: the factual pool is 2467 July draws, 1667 of them the July subset of the Phase-1 fan-out (seeds 1000-20999 (Phase-1 global unguided fan-out; July month subset)) and 800 a dedicated zero-perturbation top-up (seeds 41000-41799 (attribution_grid.BASE_SEED + pos, delta_0 scenario)), with 2500 draws at each perturbed rung. Guided probes are smaller by construction and say so wherever they appear: the odds-ratio heatwave probe is 16 draws per scale, the enrichment and physics audit 81 per scale, the tropical-cyclone ladder 40 per rung.
Guidance strengths are quoted on the core model's own scale, where the paper's setting is 1280. The event thresholds come from the emulator's own unguided climatology over the Paris box (latitude 46 to 51, longitude −2 to 6, months 6, 7, 8), fixed at a base rate of 1% on 5001 draws before any guided draw was scored against them.
Provenance
Nothing on this page is retyped. Each figure is interpolated from a committed result JSON, exported through a pass-through layer that adds a stamp and changes no value, and every stamp is listed below. ladder.json, the one shaped rather than passed through, is built from the Phase-2 and Phase-3 weights audits and carries its own coercion log for the k̂ values that are infinite. The full trust gate, both scale tables and the refusal are rendered on the rigour page.
| File | Source | Revision | Generated |
|---|---|---|---|
pricing_tc.json | results/phase35_pricing_tc.json | f3b6cb7 | 2026-07-31 |
pricing_heat.json | results/phase35_pricing_heat.json | f3b6cb7 | 2026-07-31 |
trust_gate.json | results/phase35_trust_gate.json | f3b6cb7 | 2026-07-31 |
attribution.json | results/phase35_attribution.json | f3b6cb7 | 2026-07-31 |
thresholds.json | results/phase3_thresholds.json | f3b6cb7 | 2026-07-31 |
enrichment.json | results/phase3_physics_audit_heat.json | f3b6cb7 | 2026-07-31 |
What it cost
The whole project spent 125.4 A100-hours against an operating target of 500, which is 31.3 DKRZ node-hours once the partition's node-fraction billing is applied. The expensive part was never the unguided climatology. It was the guided probes, where a single full odds-ratio draw costs 1100 s, and the point of the equal-cost comparison above is that this is the number a desk should be shown before it commissions anything. This dashboard and its figures cost no GPU time at all.
Reproducing it
The code, the committed result JSONs behind every number here and the synthetic test suite are public at https://github.com/wienkers/tailtwist. The extreme-value engine is tailspec, developed separately and used here unmodified, which is the point: an audit that shares no code with the thing it audits. cBottle and earth2studio are NVIDIA's, under Apache 2.0, and the attribution is carried in the repository.
Phase-1 Gate-1: never generalise a single pooled anchor to an area claim -- report all three anchors (nw_bahamas, ne_caribbean, n_gulf_of_mexico) and quote region_mean_return_levels (the gridded region mean, all-year rate, from phase1_tail_audit.json) for any area-level statement. region_mean_return_levels is copied verbatim as context; it is NOT September-conditioned like the anchors above (it pools the full year at the Phase-1 1460 obs/yr rate), so it is not directly comparable to the anchor blocks' return levels -- read it as the independent area-level check, not a recomputation at the same rate.