Day 2 – Climate Data

1. Day 2 Overview

African Summer School · Day 2

Climate Data: From Observations to Model Outputs

Weather and climate applications depend on many different types of data. Today you will explore where these datasets come from, what they represent, how they are stored and how they can be prepared for analysis.

Day 2 programme

Morning session

Understanding weather and climate datasets

Explore station observations, satellite products, reanalysis, operational forecasts and climate-model outputs. Compare their strengths, limitations, spatial coverage and temporal resolution.

Afternoon session

Data formats, preprocessing and Python

Learn how weather and climate information is stored in CSV, NetCDF and GRIB files. Practise inspecting variables, dimensions, units, missing values and metadata before using the data in an AI workflow.

Learning outcomes

By the end of Day 2, you should be able to:

1. Identify data sources

Distinguish station, satellite, reanalysis, forecast and climate-model datasets.

2. Compare datasets

Explain differences in coverage, resolution, frequency, uncertainty and accessibility.

3. Recognise data formats

Recognise the basic characteristics of CSV, NetCDF and GRIB files.

4. Inspect data quality

Check variables, units, coordinates, missing values, time periods and metadata.

5. Prepare data for analysis

Describe the main preprocessing steps required before training an AI or ML model.

6. Explore data with Python

Open a dataset, examine its structure and create a simple visualisation.

Today’s data journey

Every successful weather or climate AI application begins with understanding the data before selecting a model.

1. Discover Find relevant data sources
2. Understand Read variables and metadata
3. Check Assess quality and gaps
4. Prepare Clean and transform
5. Explore Analyse and visualise

Opening reflection

Think about the data you currently use in your organisation or country.

Consider:

  • Which weather or climate datasets do you use most frequently?
  • Where do these datasets come from?
  • What limitations or data gaps affect your work?
  • Which additional datasets would improve your services?