AI Risk Management12 min read

Governing a Model That Decides Who Evacuates

By Adrian Dunkley·Sep 15, 2026

Caribbean governments are beginning to buy AI-based climate and hazard intelligence. Parametric insurers already price on modelled hazard fields. Utilities are specifying design conditions from downscaled projections. Within a few years, some of these systems will be inside an evacuation decision.

That is a category of AI system the existing assurance literature handles badly. The NIST AI Risk Management Framework, ISO/IEC 42001 and most vendor due-diligence checklists were built around classifiers, recommenders and language models, where the characteristic failures are unfair treatment of a protected group, an unexplainable decision, a hallucinated fact, or a privacy breach. Those controls are worth having. They will also pass a climate emulator that is creating mass at its subdomain boundaries, and they will pass a hazard field that is confidently wrong because it faithfully reproduced a biased input.

How a Physical Model Fails

A hazard model produces fields: wind, rainfall, temperature, water depth. It fails in ways that have no analogue in a credit-scoring model.

It can violate physics undetected. Mass, energy and momentum have to be conserved. A learned model has no obligation to respect that unless you impose it, and a field that violates conservation can still score well on root-mean-square error, because the error metric is averaging over exactly the region where the violation cancels.

It can inherit a bias and look better for it. Physics-constrained downscaling models currently enforce a constraint that the average of the fine output should match the coarse input. That improves every diagnostic you would normally run. It also assumes the coarse input is an unbiased aggregate of the truth, which is true in a controlled experiment where the coarse field was made by averaging the fine one, and false in operation, where the coarse field comes from an independent global model. Global models carry documented precipitation biases over the Caribbean basin. Enforcing the constraint then drags the fine output towards the bias and makes the model more confidently wrong.

It can be right on average and wrong where people are. A domain-averaged skill score hides exactly the signal a hazard model is bought for. Skill on a steep windward slope and skill on a flat plain are different quantities, and a single number for the domain reports neither.

A Control the Checklists Do Not Have

There is a second conservation law available, and it is directly auditable.

The constraint in use today is vertical: it binds the coarsened fine field to the coarse driver, cell by cell. A second constraint is horizontal: the flux of a conserved quantity leaving one subdomain across a boundary must equal the flux entering its neighbour across that same boundary, at every timestep.

These are independent properties. A field can satisfy the first exactly at every cell while violating the second badly, because a cell mean is preserved regardless of how mass moves within or between cells.

Cross-scale and lateral conservationThree panels. The first shows the cross-scale constraint, which binds the coarsened fine field to the coarse input. The second shows the lateral constraint, which binds the flux leaving one subdomain to the flux entering its neighbour across the shared face. The third shows a field that satisfies the first exactly while violating the second, which is why the two properties are independent, so enforcing one says nothing about the other.Two conservation laws, and only one of them is being enforcedThe constraint in current physics-constrained downscaling is vertical. The one dLEGOS adds is horizontal.ACross-scale (vertical)10.0coarse driver xᴸᴿ𝒞 coarsening12987141110910fine field x̂ᴴᴿ, mean 10.0𝒞[x̂ᴴᴿ]ᵢⱼ = xᴸᴿᵢⱼHolds by construction when the coarse field is madeby coarsening the fine one. Fails when the driver isan independent model carrying its own bias.BLateral (horizontal)brick bbrick b′face fℱₑₑℱ b→b′ℱ b′→bℱ b→b′(f,t) + ℱ b′→b(f,t) = 0What leaves one brick through a face enters the nextthrough the same face. The constraint binds the finefield to itself, so no driver bias enters it.CA field that satisfies A and breaks B1288128121283.05.0both bricks average 10.0, so 𝒞[x̂] = xᴸᴿ exactlyinterface residual 2.0 kg m⁻² s⁻¹mass is created at the face and nothing detects itIndependent constraints, not nested ones.Enforcing the vertical one says nothing at all aboutthe horizontal one, and the converse also holds.
Figure 1. Two conservation laws. Cross-scale conservation binds the coarsened fine field to the coarse input, cell by cell. Lateral conservation binds the flux leaving one subdomain to the flux entering its neighbour across the shared face. Panel C constructs a field that satisfies the first exactly while violating the second by 2.0 kg m−2 s−1, which is what it means to say the two are independent.

The governance value of the second law is that it produces a number: the interface residual, in physical units, that any auditor can compute from the model's own output without needing a single observation. It is the closest thing in this domain to a control total in a financial reconciliation. It does not tell you the forecast is right, only that it is internally consistent, and a large residual is evidence of a specific defect.

The second property that makes it useful for assurance is that it binds the fine field to itself rather than to an external driver, so it does not inherit the driver's bias. That is a hypothesis, not an established result, and the paragraph below explains why the distinction matters for how you write it into a contract.

Pre-Registration as a Governance Instrument

The single most useful question to put to a vendor of a physical model is the one nobody asks: what result would make you say this model had failed?

In research this is called pre-registration. You state, before running the experiment, what the model predicts and what observation would refute it. A supplier who has done that has an answer. A supplier who has not is selling a demonstration, because a system that cannot fail cannot be validated either.

Pre-registered prediction for the bias sweepA schematic plot of conservation error against an injected bias in the driving model. The cross-scale constraint is predicted to degrade roughly in proportion to the driver bias, because it binds the fine field to the coarse field. The lateral constraint is predicted to stay flat, because it binds the fine field to itself. The plot carries no data. It states what the experiment is expected to show and what result would refute it.The claim, and the result that would kill itThis figure contains no measurements. It is the hypothesis, drawn before the runs, so it can be checked against them later.PREDICTION · NO DATA YET1×2×3×4×conservationerror0%5%10%15%20%25%30%injected bias in the driving model (% of the field mean)cross-scale constraint onlyboth constraintslateral constraint onlythis gap is the whole claimat 0% bias the two are equivalent: that is the case already testedWhat would refute itIf the teal line climbs with the red one,the lateral constraint inherits driver biasafter all and the central claim is wrong.If the two slopes are statisticallyindistinguishable across the sweep, thedistinction has no operational value evenif it is formally correct.Either outcome is reportable, and eitheranswers a question the authors of theincumbent method raised and left open.Bias is injected at known magnitude andsign, so the sweep gives a dose-responsecurve, not a single comparison.
Figure 2. The prediction, and what would refute it. Drawn before the runs, with no data in it. If the cross-scale curve climbs with driver bias and the lateral curve stays flat, the distinction between the two conservation laws has operational value and not just formal correctness. If they climb together, the claim is wrong. That is still a reportable result, because it answers a question the authors of the incumbent method raised in their own discussion and could not settle.

For a hazard model, the pre-registered test that matters most is a driver-bias sweep. Inject a bias of known magnitude and sign into the driving field and measure how the output degrades. Every operational deployment runs on a biased driver, so the degradation curve is the property you are actually buying, and a headline skill score computed on an unbiased test set tells you nothing about it.

Six Clauses for a Procurement Specification

These are drafted to be testable by a risk function without a climate physicist in the room.

  1. Native resolution over land. The supplier shall state the horizontal grid spacing of the underlying simulation over land, distinct from the delivery resolution of the product. Interpolating a coarse field onto a fine grid shall not be represented as resolution.
  2. Conservation reporting. The supplier shall report conservation error and, where the model is decomposed, interface residuals, in physical units, for every delivered field. These shall be reproducible by the buyer from delivered output.
  3. Driver-bias degradation curve. The supplier shall provide skill and conservation metrics as a function of injected driver bias across a stated range, and shall identify the bias level at which the product ceases to meet its stated tolerance.
  4. Stratified skill. Skill metrics shall be reported conditioned on elevation band and on windward against leeward aspect, not as domain means alone.
  5. Declared failure conditions. The supplier shall state, in advance, the observations that would constitute a failure of the model, and shall report against them.
  6. Compute cost per simulated year. The supplier shall state the cost in currency of producing one downscaled simulated year at the delivered resolution, so that the buyer can determine whether the ensembles required for probabilistic use are affordable within the contract.

Mapping to the Frameworks You Already Use

None of this requires a new standard. It sits inside the frameworks Caribbean institutions are already adopting.

Under NIST AI RMF, conservation residuals and interface residuals are MEASURE-function metrics for validity and reliability, and they have the useful property of being computable without ground truth. The driver-bias sweep belongs under MAP, because it characterises the operational context the system will actually run in. Declared failure conditions belong under GOVERN, as documented acceptance criteria.

Under ISO/IEC 42001, the six clauses above are straightforward AI system impact assessment inputs and performance evaluation criteria. The stratified skill requirement is the one most likely to be omitted, because domain-mean reporting is the industry default and looks complete.

Under model risk management practice of the kind Caribbean financial regulators increasingly expect, conservation residuals are the independent validation test that does not depend on the developer's own test set. That makes them unusually valuable to a second line of defence with limited technical depth.

Where This Framework Is Weak

Three limitations, stated because a governance note that claims completeness is worth less than one that does not.

Conservation is necessary and not sufficient. A field can be perfectly consistent across every boundary and still be wrong about the weather. These controls remove a failure mode; they do not certify skill, and a supplier who leans on a clean residual report as evidence of accuracy should be challenged on exactly that point.

The interface residual depends on an estimator. Computing flux across a boundary from a model's gridded output is itself approximate, and the estimator's error enters the number. A buyer should ask for the estimator to be specified, and should treat the residual as a relative indicator across model versions more confidently than as an absolute tolerance.

Observational validation in this region will stay thin. Caribbean rain-gauge networks are sparse and clustered, and gridded precipitation products over the region are dominated by station density more than by grid resolution, with the effect strongest under steep altitude gradients. That is the terrain these models exist to resolve. Physics-based and process-based checks are the substitute, and they are a genuine substitute, but they are weaker than a dense observing network would be and nobody should pretend otherwise.

Why CAIRMC Is Raising This Now

The window for writing these clauses into contracts is while the systems are being procured, not after a bad evacuation decision. The Caribbean is buying climate intelligence faster than it is building the capacity to assess it, which is the same pattern that produced the region's gaps in financial model governance a decade ago.

The research underlying the technical argument here is doctoral work at the Climate Studies Group Mona, in the Department of Physics at The University of the West Indies, Mona. It is in progress, it has no results yet, and the interface-residual claim is a hypothesis with a pre-registered experiment attached rather than a validated control. That is precisely the standard this note asks buyers to hold suppliers to, and it would be inconsistent to exempt the research from it.

The technical version, including the conservation-law argument and the experimental design, is published here.