Case study · Raymond Huang · 2026

A Google weather model on a gaming GPU

WeatherNext 2 is Google DeepMind's open-weights AI forecast model. Its official GPU path needs an 80 GB data-centre card. I changed how it executes, keeping the math identical, so the same model runs in 6.4 GiB and 7× faster. Then I built a live 10-day ensemble forecast on top of it that one RTX 5090 updates twice a day.

Open the live globe Live forecast, updated twice a day
The globe with 10 m wind from a recent run: animated flow lines, coastlines, and a tropical cyclone with its ensemble-member tracks. Explore layers, point forecasts and cyclonesExplore the full globe
The latest run, live: 10 m wind with the strongest active cyclone and every member's track. Drag to turn the globe.
34 → 6.4 GiBGPU memory for one full-resolution step
5.5 → 0.78 sper 6-hour step on an RTX 5090
0.9999999998minimum correlation with the reference at fp32, over all 101 output fields
≈ 5 minfor 8 members × 10 days, the whole live run

01The problem

WeatherNext 2 forecasts the whole atmosphere at 0.25° (about 28 km) in 6-hour steps, and every call draws a new, physically consistent ensemble member. Google publishes the weights. The code, though, is written for TPUs, and its GPU path needs about 34 GiB for a single step. In practice that means an A100-80GB or H100.

The memory goes to three places: a 24-layer transformer whose attention mask covers 32 hops of an icosahedral mesh, two graph networks that materialise per-edge tensors (the mesh-to-grid network has 3.1 million edges, 8.9 GiB per tensor), and the fragmented heap that 24 unrolled layers leave behind.

02What I changed

faster_weathernext.enable() swaps a few classes in the official modules for subclasses that compute the same function with far less memory traffic. The weights load unchanged; nothing is quantised, distilled or approximated.

  1. 1
    A fused attention kernel. A Pallas (Triton) flash-attention kernel for the 32-hop mesh mask. It visits only the 64×32 tiles that contain a valid key and never writes the attention matrix to memory. Ordering mesh nodes along a 3D Hilbert curve halves the tiles it has to visit.
  2. 2
    Blocked graph networks. Edges and grid points are processed in blocks with in-place updates, so the 8.9 GiB edge tensors never exist in full. The encoder and decoder are blocked the same way.
  3. 3
    No wasted memory. The 24 transformer layers run as a loop (one layer's buffers instead of a fragmented heap), and the attention mask exists once instead of as 24 inlined copies.
Bar charts: GPU memory for one 0.25° step, 34 GiB for the official GPU path versus 6.4 GiB; time per step on an RTX 5090, 5.47 s for the reference versus 0.78 s. Bar charts: GPU memory for one 0.25° step, 34 GiB for the official GPU path versus 6.4 GiB; time per step on an RTX 5090, 5.47 s for the reference versus 0.78 s.

The same code now runs on cards the official path cannot use:

GPUMemorySeconds per step10-day member
RTX 509032 GB0.7831 s
A10040 GB1.144 s
A10G24 GB2.51.7 min
L424 GB3.52.3 min
RTX 40608 GB5.23.5 min
T4 (no TF32)16 GB4429 min

03Same model, checked

A faster model that forecasts differently would be worthless, so equivalence came first. One step with the same inputs and noise, all 101 output fields, against a reference that is itself matched to the official GPU path on an A100-80GB:

Matmul precisionMin correlationMax relative RMS difference
fp32 (highest)0.99999999982.0 × 10⁻⁵
TF32 (GPU default)0.999994.8 × 10⁻³, the size of TF32's own rounding

In a chaotic atmosphere any rounding difference grows, so I also re-ran nor'easter hindcasts from four start times with all four checkpoints and the same noise seeds. The difference between implementations stays far below the spread between ensemble members, 181× smaller at +6 h and still 7× smaller at +150 h, and skill against ECMWF analyses is unchanged within sampling noise.

Sea-level pressure difference versus lead time: the difference between faster-weathernext and the reference stays 181 times below the member spread at 6 h, 11 times at 72 h and 7 times at 150 h. Sea-level pressure difference versus lead time: the difference between faster-weathernext and the reference stays 181 times below the member spread at 6 h, 11 times at 72 h and 7 times at 150 h.

Every number links to raw logs and scripts in the validation notes, including cross-checks on A100, A10G, L4, T4 and RTX 4060.

04The product: an ensemble you can read

Most weather globes show one forecast. An ensemble says how sure that forecast is, so the globe is built around the spread between members:

Temperature layer over Japan with a point forecast for Tokyo: temperature band across members, rain bars with chance of rain, and wind.
The globe on a phone: layer chips at the top, the western Pacific with a cyclone, and the timeline at the bottom.
Left: a Tokyo point forecast on the temperature layer. Right: the phone layout.

05How the site stays live

The site is static; the forecast is the only thing that needs a GPU. Twice a day, after ECMWF publishes the 00 and 12 UTC analyses, this pipeline runs:

  1. Initial conditionsECMWF IFS open data, fetched from Google's mirror with fallbacks, since the AWS bucket throttles fresh runs.
  2. Forecast at homeRTX 5090, 8 members × 40 steps in about 5 minutes. It waits while the machine's other jobs are running.
  3. Cloud fallbackIf the home run misses its window, a Modal watchdog runs it on an A100 (about $0.35, capped at $10 a month).
  4. EncodeEnsemble mean, spread and member range quantised into 0.5° PNG frames; cyclone tracks from every member.
  5. PublishCloudflare Pages. The browser decodes and colours the fields on the GPU with three.js.
Where a run happensTime for 8 × 10 daysCost per run
RTX 5090 at home≈ 5 min≈ $0.01 of electricity
A100 40GB on Modal (fallback)≈ 9 min≈ $0.35
Official GPU pathneeds an 80 GB card for a single step

Cloud figures are estimates from measured step times and Modal's list prices; the home figure assumes about 700 W for 6 minutes.

06Where this sits

Google runs its own interactive site, Weather Lab, with official forecasts and 64-member cyclone ensembles, and announced WeatherNext 3 in September 2026 (hourly, 5 km, driven by satellite observations, offered through Google Cloud rather than as open weights). This project does not compete with either. Its point is that the open model runs on hardware people own: the official code and checkpoints, unchanged results, on a card that fits in a desktop, for about a cent per run.

Related work I compared against or contributed to: NVIDIA's earth2studio wraps WeatherNext 2 with the official GPU attention (80 GB cards), and a Hugging Face transformers port reports about 50 GB per member. I shared the memory results with google-deepmind/weathernext#223 and on the transformers port.

07Limits, stated plainly