Differentiable Climate Emulation in JAX
Data Science Capstone - DSC 180A/B Section B03 (TA: Amirhossein Panahi, apanahi@ucsd.edu)
Introduction to Topic
The choices humanity makes in the next few decades will determine how much warmer the Earth will be by the end of the century, with implications for billions of lives and trillions of dollars in GDP. Many different emission pathways are compatible with the Paris climate agreement, and many more miss that target. Full-complexity Earth System Models can only simulate a small subset of these scenarios, so fast, accurate emulators are essential for exploring the space of possibilities and quantifying the associated risks. Our lab developed ClimateBench and more recently ClimateBench2 as a benchmark for this problem that unlock the possibility of more accurate and skillful climate models.
This project will develop direct climate emulators in JAX that target these datasets and integrate with our differentiable JAX Earth Model: the JAX-GCM (JCM) atmosphere, coupled to ocean, land, and sea-ice components through JAX-ESM (JEM). Students will explore architectures (neural operators, transformers, pattern-scaling hybrids), physical constraints (energy conservation, spatial coherence), and the use of jax.grad for gradient-based scenario analysis and calibration. The goal is a new generation of emulators that are fast, reliable, differentiable, and suitable for integration into downstream decision support.
The two core papers are:
Watson-Parris, D., Rao, Y., Olivié, D., Seland, Ø., … “ClimateBench v1.0: A benchmark for data-driven climate projections”. Journal of Advances in Modeling Earth Systems 14, e2021MS002954: https://doi.org/10.1029/2021MS002954
Davenport, E. H., Madan, J. V., … Watson-Parris, D. “JCM v1.1: a differentiable, intermediate-complexity atmospheric model”. Geoscientific Model Development 19, 6451–6466: https://doi.org/10.5194/gmd-19-6451-2026
How the year is structured
Phase I (DSC 180A, Fall) has three parts:
- Reproduce ClimateBench (weeks 1–3). Reproducing a paper’s results confirms that its findings are robust rather than the product of chance or error, and it is the fastest way to understand the methods in depth. You have AI coding assistants to help you, so this now takes about a week of focused work instead of most of the quarter. The bar for understanding what you have reproduced has not changed (see Working with AI assistants).
- Run a climate model yourself (weeks 4–7). JCM is a full atmospheric general circulation model written in Python/JAX, and its SPEEDY configuration is cheap enough to run on a single GPU (or even a laptop). You will diagnose radiative forcing with fixed SSTs, then couple it to a slab ocean with JEM to run CO₂ experiments to equilibrium and measure climate sensitivity directly. Finally, you will use
jax.gradthrough the coupled model itself. This gives you physical intuition for what an emulator is actually emulating. - Bring the two together (weeks 8–10). You will write a differentiable ClimateBench emulator in JAX, use its gradients for scenario analysis, and write a Phase II proposal.
Phase II (DSC 180B, Winter) is your own project, building on Phase I. See Phase II ideas below. For an example of what a strong project looks like, see the award-winning SeeRise project from 2025.
Compute and data
Use DataHub by default. Each of you has a GPU on DataHub, which is enough for everything in Phase I: working with the processed ClimateBench data, training emulators, and running JCM. Start there.
Use Casper only when you need the raw data it holds. Casper is the data analysis cluster at the National Center for Atmospheric Research (NCAR). We use it for the raw CMIP6 output (week 2) and ERA5 (week 4). It has a steeper learning curve, and our core hours are limited and shared across the class, so keep jobs small and ask Amirhossein or me before running anything long or GPU-heavy there. It is also a national facility with shared resources, so follow their rules and don’t leave idle jobs or Jupyter sessions running.
The data comes from the sixth Coupled Model Intercomparison Project (CMIP6), the combined effort of dozens of international modelling centres running hundreds of thousands of simulation years. The full archive (about 30 PB) is publicly available through ESGF and AWS. The processed ClimateBench dataset is on Zenodo (download it to DataHub), and the raw NorESM2 output is already on Casper.
Schedule
Click the “topic” links below for the readings, questions, and tasks for each week. Later weeks will be fleshed out as we go. Optional background papers for each part of the course are on the background reading page.
| Week | Topic | Key deliverable |
|---|---|---|
| Summer | Summer preparation | |
| 1–2 | Topic, paper, and data | Answers to the reading questions; maps and time series of the ClimateBench targets; one experiment regenerated from raw CMIP6 |
| 3 | ClimateBench in a week | Leaderboard reproduced for all 4 baselines; a 5-minute replication presentation |
| 4–5 | Running JCM: forcing, feedbacks, and climate sensitivity | JCM paper discussion; SPEEDY control climatology vs. ERA5; fixed-SST ERF vs. CO₂; slab-ocean (JEM) abrupt 2×/4×CO₂ runs, Gregory regression, and ECS |
| 6 | Differentiating a climate model | jax.grad of ERF or slab-ocean warming vs. finite differences; parameter sensitivity maps; a twin-experiment calibration |
| 7 | Emulating JCM itself | Small perturbed-parameter ensemble of slab-ocean CO₂ runs; a Gaussian process emulator of ERF, feedback, and ECS as functions of the parameters |
| 8 | A differentiable ClimateBench emulator | Best baseline ported to JAX; jax.grad of regional warming with respect to emissions; a simple inverse problem (for example, the largest emissions compatible with ΔT < 2 K) |
| 9 | Validation and Phase II proposal | Physical consistency checks on your emulator; a 1–2 page Phase II proposal |
| 10 | Wrap up and debrief | Final Phase I presentation |
Working with AI assistants
You are encouraged to use AI coding assistants (such as Claude Code) throughout this project. The JCM repository even ships a CLAUDE.md to help them. This changes how fast you can go, but not what counts as done:
- You own every line. Every team member should be able to explain any code in your repository, including why each preprocessing step is there. Expect to be asked in section.
- Verify against independent evidence. Code that runs is not a reproduction. Your numbers must match something you didn’t generate with the same tool: the published leaderboard, the Zenodo files, a finite-difference check, or a known physical value.
- Watch for the classic failure modes. Test-set leakage, silently mis-ordered dimensions (for example, assuming which vertical index is the surface), wrong area weighting, and unit errors. Assistants make these mistakes confidently.
- Keep a short decision log (
NOTES.mdin your repo). Record what you asked, what you changed, and what you checked. It will make your final report much easier to write.
Phase II ideas
These are starting points, not assignments. The strongest projects usually come from something you noticed in Phase I.
- Emulator architectures on ClimateBench/ClimateBench2. Try neural operators (FNO/SFNO), transformers, or graph networks, and compare them with the baselines under a fixed compute budget. Where do they gain skill (regional patterns, precipitation extremes, aerosol-driven responses)?
- Physically constrained emulators. Build pattern-scaling hybrids in which a JAX energy-balance model (calibrated by gradient) drives learned spatial patterns. Try enforcing a global energy budget or spatial coherence, and test whether the constraints help out-of-distribution scenarios.
- Gradient-based scenario analysis. Use
jax.gradthrough a differentiable emulator to find optimal or “Paris-compatible” emission pathways, attribute regional warming to regional emissions, or compute carbon budgets with uncertainty. - A JCM–JEM benchmark. Coupled to a slab ocean, JCM can run transient emission scenarios, and it is cheap enough to run many more of them than CMIP6 provides. Design and generate a ClimateBench-style dataset from JCM–JEM, and use it to study how emulator skill scales with the number and diversity of training scenarios, or whether pre-training on it helps emulate NorESM2.
- Calibrating JCM. Tune SPEEDY parameters against ERA5 by gradient descent. Does the calibration constrain the model’s forcing and feedbacks, or do very different sensitivities fit the present-day climate equally well?
- Learning the ocean. Every slab-ocean parameter, and the Q-flux itself, is differentiable in JEM. Train the Q-flux or mixed-layer depth by gradient descent against observed SSTs or NorESM2’s ocean heat uptake, or replace it with a small neural network. This one is for the model-development inclined; a more ambitious version compares against JEM’s experimental coupling to the Veros ocean GCM.
- Decision-support applications. Build an interactive tool on top of your emulator for a specific stakeholder question, as SeeRise did.
Section & Group Participation
Participation in the weekly discussion section is mandatory. During Phase I, each week you are responsible for doing the reading/task assigned in the schedule and submitting answers to the listed questions before discussion section begins.
Weekly assigned questions help me see how you are all doing on the project, focus your work for the week, and help you prepare for discussion. If you have questions about your work, please ask them in section, on Slack, or in office hours (I will rarely comment on your submission answers).
Office Hours
Mondays 1-2pm in HDSI room 455.
You’re also welcome to join our group meetings on Mondays at 3:30pm in Nierenberg Hall room 400 at SIO.