Microsoft Research ML pipeline forecasts space-weather risk to 66,935 US substations

Microsoft Research ML pipeline forecasts space-weather risk to 66,935 US substations

A Microsoft Research blog post, written in the first person by a summer intern, describes a machine learning system that forecasts space-weather risk across 66,935 substations in the continental United States. The aim is to give grid operators a location-specific warning 30 to 60 minutes before a specific risk appears. The post opens with the May 2024 geomagnetic storm, when utilities across North America prepared for possible impacts as auroras extended far beyond their usual range. Extreme space-weather events can induce currents in transmission networks that damage equipment and raise operational risk, and the challenge is not only knowing a storm is coming but estimating when and where its effects could be most severe.

The pipeline has three stages. First, solar-wind measurements from the L1 Lagrange point are used to forecast two indices, Auroral Electrojet (AE) and Disturbance Storm Time (Dst), while geological conductivity and location features are assembled for each substation. Second, a gradient-boosting model combines the forecasted and location-specific inputs to estimate dB/dt, the rate of magnetic-field change associated with geomagnetically induced current (GIC) risk. Third, the predictions are converted into location-specific risk estimates and aggregated into a continental risk assessment. The author says a system of 50 AI agents helped explore features, validation strategies and model configurations. Only public data was used: NASA OMNI and NASA-aggregated Kyoto World Data Center data, INTERMAGNET and U.S. Geological Survey magnetometer observations, and GridSFM-derived grid data.

On evaluation, the AE predictor targets rare, high-intensity geomagnetic activity. Over the 2020-2026 evaluation period it produced forecasts spanning nearly the full observed AE range and outperformed several empirical solar-wind-based approaches. The Dst predictor adds a signal for large-scale storm strength. During the most geomagnetically active periods of 2020-2026, the machine learning model beat the Burton equation on 62.2% of individual hours, produced a substantially wider prediction range than Burton-style approaches, and, combined with AE forecasts, improved severe-event detection in the end-to-end system by 1.2 percentage points.

The GIC risk stage was judged differently. The author says no equivalent widely deployed operational system offers a direct industry benchmark, so the model was compared with simple linear regression. It detected 76.5% of major events (at least 10 nT/min), 81.2% of severe events (at least 20 nT/min) and 64.1% of extreme events (at least 50 nT/min). The post's summary box puts the overall figure at nearly 80% of major events. False-alarm rates rose with storm severity, which the author describes as the trade-off between missed events and cautious alerts. Detection varied by latitude and was highest at northern stations, where geomagnetic activity is strongest.

Rather than one alert for the whole country, the system combines storm conditions with each substation's latitude and geological factor, giving continuous risk estimates that separate lower-risk sites from areas where resistive geology can amplify ground-level effects. A demonstration map for a representative major-storm scenario is shown, and the author stresses it is not a record of a live operational event. Estimates for all 66,935 substations took approximately 333 milliseconds during measured inference, so many scenarios can be run quickly.

Looking ahead, the author says earlier, location-specific information could help utilities prioritize engineering review and consider protective steps such as adjusting reactive-power reserves or temporarily reconfiguring parts of the network. Further validation with utilities and operational data would be needed before the system could be used in grid operations. Listed next steps: longer forecast horizons using temporal-transformer approaches, adaptation to other regions, integration with existing grid decision workflows, and transformer-level risk instead of substation-level estimates. The work is tied to other Microsoft Research efforts on AI for power systems, including GridSFM, which applies deep learning to AC optimal power flow.

Key facts

  • The pipeline produces location-specific risk estimates for 66,935 substations in the continental United States, 30 to 60 minutes before a specific risk appears.
  • It forecasts AE and Dst indices from L1 solar-wind data, then a gradient-boosting model estimates dB/dt using each substation's latitude and geology.
  • Detection rates at the GIC risk stage: 76.5% for major events (at least 10 nT/min), 81.2% for severe (at least 20 nT/min) and 64.1% for extreme (at least 50 nT/min); false alarms rose with severity.
  • The ML Dst model beat the Burton equation on 62.2% of individual hours in the most active periods of 2020-2026, and added 1.2 percentage points to severe-event detection when combined with AE forecasts.
  • The author says further validation with utilities and operational data is needed before grid use; inference for all substations took about 333 milliseconds.

Why it matters

Modern society depends on reliable electric power, and extreme space-weather events can induce currents in transmission networks that damage equipment. The post argues that the hard part is not knowing a storm is approaching but estimating when and where its effects could be most severe, with enough warning to act. The system tries to turn a broad space-weather warning into a per-substation view, so planners could see which locations may warrant closer analysis. The author frames it as a demonstration of how physics-grounded machine learning could support more specific and timely assessment of exposure.

Who it affects

The intended users are grid operators, planners and utilities, who could use earlier, location-specific information to prioritize engineering review and consider protective actions such as adjusting reactive-power reserves or temporarily reconfiguring parts of the network. The estimates cover substations in the continental United States. The post lists international adaptation, with different geology and grid topologies, as a future step.

How to use it

This is a research prototype described in a blog post, not a product. The source does not state that any utility uses it. The author says further validation with utilities and operational data is needed before grid operations. The data inputs are public: NASA OMNI, Kyoto World Data Center data aggregated by NASA, INTERMAGNET and USGS magnetometer observations, and GridSFM-derived grid data. Planned next steps include longer forecast horizons with temporal transformers, other regions, integration with existing engineering-review workflows before any higher automation, and transformer-level risk.

How solid is it

The post is a first-person account by a summer intern at Microsoft Research, who thanks mentors Weiwei Yang and Spencer Fowers. The numbers are the author's own evaluation over 2020-2026. The AE model is said to outperform several empirical solar-wind-based approaches, but no numeric results are given for that comparison. For the GIC risk stage there is no directly comparable operational benchmark, so the baseline is simple linear regression, with no baseline numbers given. The headline summary says nearly 80% of major events were detected, while the detailed figure for major events at 10 nT/min or more is 76.5%; the post does not reconcile the two. The 62.2% result covers only individual hours in the most geomagnetically active periods, measured against the Burton equation.

Risks and caveats

False-alarm rates increased with storm severity, and no figures are given. Extreme events (at least 50 nT/min) had the lowest detection rate, at 64.1%. Performance varied by latitude, best at northern stations. The continental map is a scenario demonstration, not a record of a live event. The warning window is limited to 30-60 minutes. The author says validation with utilities and operational data is still required before use in grid operations.

“The map is a demonstration of the model’s continental-scale output, not a record of a live operational event.”

— Author of the Microsoft Research blog post