WeatherNext 2 is Google DeepMind's open-weights AI forecast model. Its official GPU path needs an 80 GB data-centre card. I changed how it executes, keeping the math identical, so the same model runs in 6.4 GiB and 7× faster. Then I built a live 10-day ensemble forecast on top of it that one RTX 5090 updates twice a day.
Explore layers, point forecasts and cyclonesExplore the full globe
WeatherNext 2 forecasts the whole atmosphere at 0.25° (about 28 km) in 6-hour steps, and every call draws a new, physically consistent ensemble member. Google publishes the weights. The code, though, is written for TPUs, and its GPU path needs about 34 GiB for a single step. In practice that means an A100-80GB or H100.
The memory goes to three places: a 24-layer transformer whose attention mask covers 32 hops of an icosahedral mesh, two graph networks that materialise per-edge tensors (the mesh-to-grid network has 3.1 million edges, 8.9 GiB per tensor), and the fragmented heap that 24 unrolled layers leave behind.
faster_weathernext.enable() swaps a few classes in the official modules for subclasses that compute the same function with far less memory traffic. The weights load unchanged; nothing is quantised, distilled or approximated.
The same code now runs on cards the official path cannot use:
| GPU | Memory | Seconds per step | 10-day member |
|---|---|---|---|
| RTX 5090 | 32 GB | 0.78 | 31 s |
| A100 | 40 GB | 1.1 | 44 s |
| A10G | 24 GB | 2.5 | 1.7 min |
| L4 | 24 GB | 3.5 | 2.3 min |
| RTX 4060 | 8 GB | 5.2 | 3.5 min |
| T4 (no TF32) | 16 GB | 44 | 29 min |
A faster model that forecasts differently would be worthless, so equivalence came first. One step with the same inputs and noise, all 101 output fields, against a reference that is itself matched to the official GPU path on an A100-80GB:
| Matmul precision | Min correlation | Max relative RMS difference |
|---|---|---|
| fp32 (highest) | 0.9999999998 | 2.0 × 10⁻⁵ |
| TF32 (GPU default) | 0.99999 | 4.8 × 10⁻³, the size of TF32's own rounding |
In a chaotic atmosphere any rounding difference grows, so I also re-ran nor'easter hindcasts from four start times with all four checkpoints and the same noise seeds. The difference between implementations stays far below the spread between ensemble members, 181× smaller at +6 h and still 7× smaller at +150 h, and skill against ECMWF analyses is unchanged within sampling noise.
Every number links to raw logs and scripts in the validation notes, including cross-checks on A100, A10G, L4, T4 and RTX 4060.
Most weather globes show one forecast. An ensemble says how sure that forecast is, so the globe is built around the spread between members:


The site is static; the forecast is the only thing that needs a GPU. Twice a day, after ECMWF publishes the 00 and 12 UTC analyses, this pipeline runs:
| Where a run happens | Time for 8 × 10 days | Cost per run |
|---|---|---|
| RTX 5090 at home | ≈ 5 min | ≈ $0.01 of electricity |
| A100 40GB on Modal (fallback) | ≈ 9 min | ≈ $0.35 |
| Official GPU path | needs an 80 GB card for a single step | |
Cloud figures are estimates from measured step times and Modal's list prices; the home figure assumes about 700 W for 6 minutes.
Google runs its own interactive site, Weather Lab, with official forecasts and 64-member cyclone ensembles, and announced WeatherNext 3 in September 2026 (hourly, 5 km, driven by satellite observations, offered through Google Cloud rather than as open weights). This project does not compete with either. Its point is that the open model runs on hardware people own: the official code and checkpoints, unchanged results, on a card that fits in a desktop, for about a cent per run.
Related work I compared against or contributed to: NVIDIA's earth2studio wraps WeatherNext 2 with the official GPU attention (80 GB cards), and a Hugging Face transformers port reports about 50 GB per member. I shared the memory results with google-deepmind/weathernext#223 and on the transformers port.