ReRoute lets climate emulators answer what-if questions without intervention data

The paper starts from a familiar problem in scientific machine learning. Many questions ask about things that were never observed: what if the conditions, interventions or history had been different? The authors say models can predict accurately on observed data yet fail on such what-if queries when correlated inputs are varied independently. The usual remedy is to add controlled simulation data in which those factors are explicitly disentangled. The authors point out that this requires access to a simulator, can be computationally expensive, and inherits the simulator's modeling assumptions.

Their answer is ReRoute, a framework for targeted scientific what-if prediction. It combines factual data with partial mechanistic knowledge and does not need controlled intervention data for adaptation. The method has three steps. It fixes the queried input of a pretrained backbone to a reference value. It then reintroduces that input's variation through a known mechanistic pathway. Finally it fine-tunes on the original factual data, leaving downstream effects to the learned dynamics.

The authors provide a causal identification result for this construction under explicit structural assumptions, and the core argument is machine-checked in Lean.

The evaluation has three stages. First, in a controlled advection-diffusion system where exact responses are available, ReRoute achieves highly accurate counterfactual predictions. Second, on state-of-the-art climate emulation with held-out coupled-climate interventions, ReRoute reduces aggregate climate error by 18.2-31.8% under severe CO2 distribution shifts. It preserves skill under standard conditions, and it does so at a small fraction of the cost of retraining on additional controlled simulations. That comparison does not even count the substantial expense of generating such simulation data. Third, on an emulator trained from historical ERA5 reanalysis, where no counterfactual reference exists, ReRoute preserves substantially more of the surface warming implied by the observed boundary conditions under a fixed-CO2 counterfactual.

Key facts

  • ReRoute is a framework for what-if (counterfactual) prediction that combines factual data with partial mechanistic knowledge and needs no controlled intervention data for adaptation.
  • Method: fix the queried input of a pretrained backbone to a reference value, reintroduce its variation through a known mechanistic pathway, and fine-tune on the original factual data.
  • The authors give a causal identification result under explicit structural assumptions, with the core argument machine-checked in Lean.
  • On held-out coupled-climate interventions, aggregate climate error falls by 18.2-31.8% under severe CO2 distribution shifts, with skill preserved under standard conditions.
  • On an ERA5-trained emulator with no counterfactual reference, ReRoute preserves substantially more of the surface warming implied by the observed boundary conditions under a fixed-CO2 counterfactual.

Why it matters

Emulators are fast stand-ins for expensive physical simulators, but they are often asked about scenarios that the data never covered, such as a different CO2 pathway. The paper's central point is that accuracy on observed data does not guarantee accuracy on these what-if queries when correlated inputs are varied independently. The standard fix, generating extra controlled simulations, needs a simulator and inherits its assumptions. ReRoute tries to skip that step by using what is already known about the mechanism, which is the novelty claimed here.

Who it affects

The framing is aimed at people building and using scientific emulators, with climate emulation as the main test case. The abstract also frames the problem more broadly, as scientific questions about conditions, interventions or history that were never observed. Teams that cannot afford to generate large sets of controlled simulations are the ones the cost argument speaks to most directly.

How to use it

The abstract describes the recipe at a high level: take a pretrained backbone, fix the input you want to query to a reference value, feed its variation back in through a known mechanistic pathway, and fine-tune on the original factual data. That means a practitioner needs a pretrained model and some partial mechanistic knowledge of how the queried input acts. The abstract does not say whether the code or data are released.

How solid is it

This is a summary of a preprint on arXiv, and the figures below come from the authors' own abstract. The causal identification result is backed by a core argument that is machine-checked in Lean, which is a strong formal check, though the abstract does not state which structural assumptions underlie it or what part of the Lean check goes beyond the core argument. The headline number is an 18.2-31.8% reduction in aggregate climate error on held-out coupled-climate interventions under severe CO2 distribution shifts. The abstract does not state the baseline it is measured against. It also gives no numeric result for the advection-diffusion system or for the ERA5 surface-warming comparison, and the cost advantage is stated only as 'a small fraction' of retraining.

Risks and caveats

The guarantee holds under explicit structural assumptions, so the method is only as trustworthy as those assumptions are for a given system. It also relies on a known mechanistic pathway for the queried input, so it targets specific what-if queries rather than arbitrary ones. The ERA5 test has no counterfactual reference, which means that result shows a preservation of warming rather than a score against ground truth. The abstract does not name the climate emulator or the pretrained backbone used, so the results cannot be tied to a specific model from this text alone.

“Models can predict accurately on observed data yet fail on such what-if queries when correlated inputs are varied independently.”

— Abstract of the ReRoute paper, arXiv 2610.02252