Skip to content

Python API

gathering provides a fluent, type-safe API designed for interactive environments (Jupyter, Colab, VSCode) and scripting.

Task Builders

The Task class provides typed builder methods for every registered source with IDE autocompletion:

from gathering import Task

# Ground-based sources
t1 = Task.aeronet(
    station="Granada",
    product="AOD15",
    start_date="2023-07-01",
    end_date="2023-07-05",
)

t2 = Task.actris_cloudnet(
    station="hyytiala",
    date="2021-01-01",
    product=["mwr", "radar"],
)

t3 = Task.pandonia(
    station="Thessaloniki",
    start_date="2023-05-01",
    end_date="2023-05-21",
    product=["fnvh3"],
)

# Satellite sources
t4 = Task.earthcare(
    product_types=["ATL_NOM_1B"],
    collections=["EarthCAREL1InstChecked"],
    start_time="2024-07-31T13:00:00Z",
    end_time="2024-07-31T14:00:00Z",
)

t5 = Task.pace(
    instrument="OCI",
    product_type="L1C",
    start_time="2024-06-01T10:00:00Z",
    end_time="2024-06-01T14:00:00Z",
    grid={"site": "Granada", "size_yx": [100, 100], "resolution_m": 1000.0},
)

t6 = Task.modis(
    satellite="Terra",
    product_type="MOD04_L2",
    start_time="2024-07-01T10:00:00Z",
    end_time="2024-07-01T14:00:00Z",
    grid={"site": "Granada", "size_yx": [100, 100], "resolution_m": 10000.0},
)

Tip

examples/all_options.py lists every parameter each gatherer accepts, with its default and valid values.

Pipeline

The Pipeline class orchestrates multiple tasks:

from gathering import Task, Pipeline

pipeline = Pipeline([t1, t2, t3, t4])

# Preview in Jupyter (renders rich HTML table)
pipeline

Sequential Execution

# Run with interactive tqdm progress bar
results = pipeline.run(show_progress=True)

Parallel Execution & Concurrency Levels

Gathering provides two distinct levels of concurrency that can be combined:

  1. Task-Level Concurrency (Pipeline): Run multiple independent source gatherers in parallel threads:

    # Execute up to 4 gatherer tasks concurrently
    results = pipeline.run(parallel=True, max_workers=4)
    

  2. Granule-Level Concurrency (Gatherer): Download multiple granules or product chunks concurrently within a single source task:

    # Download 4 granules simultaneously for MODIS and CDSE tasks
    task_modis = Task.modis(
        satellite="Terra",
        product_type="MOD04_L2",
        start_time="2024-07-01T00:00:00Z",
        end_time="2024-07-05T23:59:59Z",
        max_workers=4,
    )
    

Accessing Results

# List of downloaded file paths
print(results.files)

# Convert to pandas DataFrame
df_summary = results.to_dataframe()

Single Task Execution

You can also run a single task directly:

from gathering import Task

task = Task.aeronet(
    station="Granada",
    product="AOD15",
    start_date="2023-07-01",
    end_date="2023-07-05",
)

result = task.run()
print(result.files)

YAML Execution from Python

from gathering import run_from_yaml

# Execute a YAML configuration file
results = run_from_yaml("examples/multi_source_campaign_pipeline.yaml")

Shared Vocabulary

Ground-based sources (AERONET, Pandonia, ACTRIS Cloudnet, ACTRIS Ares) use unified parameter names. Deprecated aliases are accepted with a warning:

Concept Canonical Name Deprecated Aliases
Measurement site station site, site_id, stations
Product / data type product data_type, fcodes, kind
Range start start_date date_from, from_date
Range end end_date date_to, to_date

Deprecation Policy

Deprecated aliases are supported throughout the 0.x series and scheduled for removal in 1.0.