
Governing a Model That Decides Who Evacuates
Caribbean governments are beginning to buy AI-based climate and hazard intelligence. Parametric insurers already price on modelled hazard fields. Utilities are specifying design conditions from downscaled projections. Within a few years, some of these systems will be inside an evacuation decision.
That is a category of AI system the existing assurance literature handles badly. The NIST AI Risk Management Framework, ISO/IEC 42001 and most vendor due-diligence checklists were built around classifiers, recommenders and language models, where the characteristic failures are unfair treatment of a protected group, an unexplainable decision, a hallucinated fact, or a privacy breach. Those controls are worth having. They will also pass a climate emulator that is creating mass at its subdomain boundaries, and they will pass a hazard field that is confidently wrong because it faithfully reproduced a biased input.
How a Physical Model Fails
A hazard model produces fields: wind, rainfall, temperature, water depth. It fails in ways that have no analogue in a credit-scoring model.
It can violate physics undetected. Mass, energy and momentum have to be conserved. A learned model has no obligation to respect that unless you impose it, and a field that violates conservation can still score well on root-mean-square error, because the error metric is averaging over exactly the region where the violation cancels.
It can inherit a bias and look better for it. Physics-constrained downscaling models currently enforce a constraint that the average of the fine output should match the coarse input. That improves every diagnostic you would normally run. It also assumes the coarse input is an unbiased aggregate of the truth, which is true in a controlled experiment where the coarse field was made by averaging the fine one, and false in operation, where the coarse field comes from an independent global model. Global models carry documented precipitation biases over the Caribbean basin. Enforcing the constraint then drags the fine output towards the bias and makes the model more confidently wrong.
It can be right on average and wrong where people are. A domain-averaged skill score hides exactly the signal a hazard model is bought for. Skill on a steep windward slope and skill on a flat plain are different quantities, and a single number for the domain reports neither.
A Control the Checklists Do Not Have
There is a second conservation law available, and it is directly auditable.
The constraint in use today is vertical: it binds the coarsened fine field to the coarse driver, cell by cell. A second constraint is horizontal: the flux of a conserved quantity leaving one subdomain across a boundary must equal the flux entering its neighbour across that same boundary, at every timestep.
These are independent properties. A field can satisfy the first exactly at every cell while violating the second badly, because a cell mean is preserved regardless of how mass moves within or between cells.
The governance value of the second law is that it produces a number: the interface residual, in physical units, that any auditor can compute from the model's own output without needing a single observation. It is the closest thing in this domain to a control total in a financial reconciliation. It does not tell you the forecast is right, only that it is internally consistent, and a large residual is evidence of a specific defect.
The second property that makes it useful for assurance is that it binds the fine field to itself rather than to an external driver, so it does not inherit the driver's bias. That is a hypothesis, not an established result, and the paragraph below explains why the distinction matters for how you write it into a contract.
Pre-Registration as a Governance Instrument
The single most useful question to put to a vendor of a physical model is the one nobody asks: what result would make you say this model had failed?
In research this is called pre-registration. You state, before running the experiment, what the model predicts and what observation would refute it. A supplier who has done that has an answer. A supplier who has not is selling a demonstration, because a system that cannot fail cannot be validated either.
For a hazard model, the pre-registered test that matters most is a driver-bias sweep. Inject a bias of known magnitude and sign into the driving field and measure how the output degrades. Every operational deployment runs on a biased driver, so the degradation curve is the property you are actually buying, and a headline skill score computed on an unbiased test set tells you nothing about it.
Six Clauses for a Procurement Specification
These are drafted to be testable by a risk function without a climate physicist in the room.
- Native resolution over land. The supplier shall state the horizontal grid spacing of the underlying simulation over land, distinct from the delivery resolution of the product. Interpolating a coarse field onto a fine grid shall not be represented as resolution.
- Conservation reporting. The supplier shall report conservation error and, where the model is decomposed, interface residuals, in physical units, for every delivered field. These shall be reproducible by the buyer from delivered output.
- Driver-bias degradation curve. The supplier shall provide skill and conservation metrics as a function of injected driver bias across a stated range, and shall identify the bias level at which the product ceases to meet its stated tolerance.
- Stratified skill. Skill metrics shall be reported conditioned on elevation band and on windward against leeward aspect, not as domain means alone.
- Declared failure conditions. The supplier shall state, in advance, the observations that would constitute a failure of the model, and shall report against them.
- Compute cost per simulated year. The supplier shall state the cost in currency of producing one downscaled simulated year at the delivered resolution, so that the buyer can determine whether the ensembles required for probabilistic use are affordable within the contract.
Mapping to the Frameworks You Already Use
None of this requires a new standard. It sits inside the frameworks Caribbean institutions are already adopting.
Under NIST AI RMF, conservation residuals and interface residuals are MEASURE-function metrics for validity and reliability, and they have the useful property of being computable without ground truth. The driver-bias sweep belongs under MAP, because it characterises the operational context the system will actually run in. Declared failure conditions belong under GOVERN, as documented acceptance criteria.
Under ISO/IEC 42001, the six clauses above are straightforward AI system impact assessment inputs and performance evaluation criteria. The stratified skill requirement is the one most likely to be omitted, because domain-mean reporting is the industry default and looks complete.
Under model risk management practice of the kind Caribbean financial regulators increasingly expect, conservation residuals are the independent validation test that does not depend on the developer's own test set. That makes them unusually valuable to a second line of defence with limited technical depth.
Where This Framework Is Weak
Three limitations, stated because a governance note that claims completeness is worth less than one that does not.
Conservation is necessary and not sufficient. A field can be perfectly consistent across every boundary and still be wrong about the weather. These controls remove a failure mode; they do not certify skill, and a supplier who leans on a clean residual report as evidence of accuracy should be challenged on exactly that point.
The interface residual depends on an estimator. Computing flux across a boundary from a model's gridded output is itself approximate, and the estimator's error enters the number. A buyer should ask for the estimator to be specified, and should treat the residual as a relative indicator across model versions more confidently than as an absolute tolerance.
Observational validation in this region will stay thin. Caribbean rain-gauge networks are sparse and clustered, and gridded precipitation products over the region are dominated by station density more than by grid resolution, with the effect strongest under steep altitude gradients. That is the terrain these models exist to resolve. Physics-based and process-based checks are the substitute, and they are a genuine substitute, but they are weaker than a dense observing network would be and nobody should pretend otherwise.
Why CAIRMC Is Raising This Now
The window for writing these clauses into contracts is while the systems are being procured, not after a bad evacuation decision. The Caribbean is buying climate intelligence faster than it is building the capacity to assess it, which is the same pattern that produced the region's gaps in financial model governance a decade ago.
The research underlying the technical argument here is doctoral work at the Climate Studies Group Mona, in the Department of Physics at The University of the West Indies, Mona. It is in progress, it has no results yet, and the interface-residual claim is a hypothesis with a pre-registered experiment attached rather than a validated control. That is precisely the standard this note asks buyers to hold suppliers to, and it would be inconsistent to exempt the research from it.
The technical version, including the conservation-law argument and the experimental design, is published here.