Buckets:

18.2 TB
9,012 files
Updated 2 days ago
Name
Size
2026-09-18-downscaling-data-X-SHiELD-AMIP-for-hf-2013-2023
hirov1
hirov2
xshield
LICENSE18.7 kB
xet
README.md4.86 kB
xet
README.md

X-SHiELD 3 km 10-year dataset and HiRO-ACE outputs

This bucket contains data from GFDL's X-SHiELD, a global 3 km storm-resolving atmospheric model, used to train 100 km climate model emulators and 3 km downscaling models, and generated outputs from the HiRO-ACE, our 3 km emulation framework, for the precipitation only HiROv1, and the manuscript in preparation for a wind/pressure/precipitation HiROv2 model.

The X-SHiELD simulation covers an 11-year run in AMIP mode (forced by observed sea surface temperatures). We omit the first year of data as spin up and generally use the 2014-2022 period for training, leaving the 2023 year out for independent assessment. The HiRO-ACE generations provided in this bucket also use the year 2023 forcing, unless noted otherwise.

The bucket contains two kinds of downscaled data:

  • Perfect prediction (pp / perfect): HiRO applied to X-SHiELD output coarsened to ~100 km. The coarse inputs follow the same weather trajectory as the 3 km X-SHiELD reference, so the downscaled output can be compared with the reference timestep by timestep.
  • HiRO-ACE (ace2s / hiro-ace): HiRO applied to output from ACE2S, a ~100 km atmospheric emulator, run over 2023 with the same sea surface temperature forcing. The ACE2S run is an independent realization: its forcing matches X-SHiELD but its weather does not. Compare it with the reference statistically (for example climatologies, distributions and extremes), not timestep by timestep.

We include generated global output examples from two HiRO models:

  • HiRO-ACEv1: ACE2S pretrained on ERA5 and finetuned on X-SHiELD with only precipitation downscaling to 3 km (arxiv)
  • HiRO-ACEv2: ACE2S-XSHiELD pretrained on SHiELD+ (arxiv) and finetuned X-SHiELD with winds/sea-level pressure/precipitation downscaling to 3 km ((arxiv)[])
Store Description Time range
hirov2/2026-08-05-global-hirov2-ace2s-2023.zarr HiRO-ACE v2: global 3 km downscaling of ACE2S 100 km output. 2023
hirov2/2026-08-03-global-hirov2-pp-2023.zarr HiRO v2 perfect prediction: global 3 km downscaling of coarsened X-SHiELD. 2023
hirov2/output_6hourly_predictions_ic0000_2023-onward-rechunked.zarr ACE2S 100 km 6-hourly output that is the input to HiRO-ACE v2. 2023
hirov1/global_hiro_ace_2023.zarr HiRO-ACE v1: global 3 km downscaling of ACE2S 100 km output. 2023 (TODO)
hirov1/output_6hourly_predictions_ic0000.zarr ACE2S 100 km 6-hourly output that is the input to HiRO-ACE v1. 2023
hirov1/global_hiro_perfect_2023.zarr HiRO v1 perfect prediction: global 3 km downscaling of coarsened X-SHiELD. 2023
xshield/xshield_100km_2023.zarr X-SHiELD coarsened to ~100 km; the input to the perfect-prediction runs. 2023
xshield/2026-09-29-x-shield-3km-2023-only.zarr X-SHiELD 3 km reference. 2023
2026-09-18-downscaling-data-X-SHiELD-AMIP-for-hf-2013-2023 X-SHiELD 3 km surface fields (PRATEsfc, PRMSL, eastward_wind_at_ten_meters, northward_wind_at_ten_meters, HGTsfc, land fraction) 2013-2023

TODO: link to configs/code (e.g. the ai2cm/ace repository) and the interactive demo.

Opening data directly with xarray

With the huggingface_hub Python package installed in your environment, you can open and interact with these datasets directly with xarray. There are no egress fees and therefore no credentials are required. Opening a dataset is as simple as:

>>> import xarray as xr
>>> STORE = "hf://buckets/allenai/ai2cm-downscaling/xshield/xshield_100km_2023.zarr"
>>> ds = xr.open_zarr(STORE, chunks={"time": 400})

This is an example of the 100 km dataset, which is sharded/chunked over the time dimension. Here chunks={"time": 400} is specified to override the native inner chunk size along the "time" dimension, which is 1 for optimal performance during training. For analyses that depend on aggregations across the "time" dimension it is significantly more performant to use the shard size, in this case 400, as the dask chunk size. If your use-case more resembles the training regime, then you should likely omit it. For 3 km data, shards are large (again to reduce total number of file objects), but it is also chunked in space to retain manageable chunk sizes. To investigate the underlying store chunking/sharding information:

# example looking at a single variable from the ds loaded above
>>> ds.PRATEsfc.encoding

License

This dataset is licensed under CC BY 4.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.

Total size
18.2 TB
Files
9,012
Last updated
Oct 8
Pre-warmed CDN
US EU US EU

Contributors