Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| 2026-09-18-downscaling-data-X-SHiELD-AMIP-for-hf-2013-2023 | 3,231 items | ||
| hirov1 | 917 items | ||
| hirov2 | 3,353 items | ||
| xshield | 1,509 items | ||
| LICENSE | 18.7 kB xet | ddd2a7aa | |
| README.md | 4.86 kB xet | a3464043 |
X-SHiELD 3 km 10-year dataset and HiRO-ACE outputs
This bucket contains data from GFDL's X-SHiELD, a global 3 km storm-resolving atmospheric model, used to train 100 km climate model emulators and 3 km downscaling models, and generated outputs from the HiRO-ACE, our 3 km emulation framework, for the precipitation only HiROv1, and the manuscript in preparation for a wind/pressure/precipitation HiROv2 model.
The X-SHiELD simulation covers an 11-year run in AMIP mode (forced by observed sea surface temperatures). We omit the first year of data as spin up and generally use the 2014-2022 period for training, leaving the 2023 year out for independent assessment. The HiRO-ACE generations provided in this bucket also use the year 2023 forcing, unless noted otherwise.
The bucket contains two kinds of downscaled data:
- Perfect prediction (
pp/perfect): HiRO applied to X-SHiELD output coarsened to ~100 km. The coarse inputs follow the same weather trajectory as the 3 km X-SHiELD reference, so the downscaled output can be compared with the reference timestep by timestep. - HiRO-ACE (
ace2s/hiro-ace): HiRO applied to output from ACE2S, a ~100 km atmospheric emulator, run over 2023 with the same sea surface temperature forcing. The ACE2S run is an independent realization: its forcing matches X-SHiELD but its weather does not. Compare it with the reference statistically (for example climatologies, distributions and extremes), not timestep by timestep.
We include generated global output examples from two HiRO models:
- HiRO-ACEv1: ACE2S pretrained on ERA5 and finetuned on X-SHiELD with only precipitation downscaling to 3 km (arxiv)
- HiRO-ACEv2: ACE2S-XSHiELD pretrained on SHiELD+ (arxiv) and finetuned X-SHiELD with winds/sea-level pressure/precipitation downscaling to 3 km ((arxiv)[])
| Store | Description | Time range |
|---|---|---|
hirov2/2026-08-05-global-hirov2-ace2s-2023.zarr |
HiRO-ACE v2: global 3 km downscaling of ACE2S 100 km output. | 2023 |
hirov2/2026-08-03-global-hirov2-pp-2023.zarr |
HiRO v2 perfect prediction: global 3 km downscaling of coarsened X-SHiELD. | 2023 |
hirov2/output_6hourly_predictions_ic0000_2023-onward-rechunked.zarr |
ACE2S 100 km 6-hourly output that is the input to HiRO-ACE v2. | 2023 |
hirov1/global_hiro_ace_2023.zarr |
HiRO-ACE v1: global 3 km downscaling of ACE2S 100 km output. | 2023 (TODO) |
hirov1/output_6hourly_predictions_ic0000.zarr |
ACE2S 100 km 6-hourly output that is the input to HiRO-ACE v1. | 2023 |
hirov1/global_hiro_perfect_2023.zarr |
HiRO v1 perfect prediction: global 3 km downscaling of coarsened X-SHiELD. | 2023 |
xshield/xshield_100km_2023.zarr |
X-SHiELD coarsened to ~100 km; the input to the perfect-prediction runs. | 2023 |
xshield/2026-09-29-x-shield-3km-2023-only.zarr |
X-SHiELD 3 km reference. | 2023 |
2026-09-18-downscaling-data-X-SHiELD-AMIP-for-hf-2013-2023 |
X-SHiELD 3 km surface fields (PRATEsfc, PRMSL, eastward_wind_at_ten_meters, northward_wind_at_ten_meters, HGTsfc, land fraction) | 2013-2023 |
TODO: link to configs/code (e.g. the ai2cm/ace repository) and the interactive demo.
Opening data directly with xarray
With the
huggingface_hub Python package
installed in your environment, you can open and interact with these datasets
directly with xarray. There are no egress fees and therefore no credentials
are required. Opening a dataset is as simple as:
>>> import xarray as xr
>>> STORE = "hf://buckets/allenai/ai2cm-downscaling/xshield/xshield_100km_2023.zarr"
>>> ds = xr.open_zarr(STORE, chunks={"time": 400})
This is an example of the 100 km dataset, which is sharded/chunked over the time dimension. Here chunks={"time": 400} is specified to override the native inner chunk size along the "time" dimension, which is 1 for optimal performance during training. For analyses that depend on aggregations across the "time" dimension it is significantly more performant to use the shard size, in this case 400, as the dask chunk size. If your use-case more resembles the training regime, then you should likely omit it. For 3 km data, shards are large (again to reduce total number of file objects), but it is also chunked in space to retain manageable chunk sizes. To investigate the underlying store chunking/sharding information:
# example looking at a single variable from the ds loaded above
>>> ds.PRATEsfc.encoding
License
This dataset is licensed under CC BY 4.0. It is intended for research and educational use in accordance with Ai2's Responsible Use Guidelines.
- Total size
- 18.2 TB
- Files
- 9,012
- Last updated
- Oct 8
- Pre-warmed CDN
- US EU US EU