Python API¶
gathering provides a fluent, type-safe API designed for interactive environments (Jupyter, Colab, VSCode) and scripting.
Task Builders¶
The Task class provides typed builder methods for every registered source with IDE autocompletion:
from gathering import Task
# Ground-based sources
t1 = Task.aeronet(
station="Granada",
product="AOD15",
start_date="2023-07-01",
end_date="2023-07-05",
)
t2 = Task.actris_cloudnet(
station="hyytiala",
date="2021-01-01",
product=["mwr", "radar"],
)
t3 = Task.pandonia(
station="Thessaloniki",
start_date="2023-05-01",
end_date="2023-05-21",
product=["fnvh3"],
)
# Satellite sources
t4 = Task.earthcare(
product_types=["ATL_NOM_1B"],
collections=["EarthCAREL1InstChecked"],
start_time="2024-07-31T13:00:00Z",
end_time="2024-07-31T14:00:00Z",
)
t5 = Task.pace(
instrument="OCI",
product_type="L1C",
start_time="2024-06-01T10:00:00Z",
end_time="2024-06-01T14:00:00Z",
grid={"site": "Granada", "size_yx": [100, 100], "resolution_m": 1000.0},
)
t6 = Task.modis(
satellite="Terra",
product_type="MOD04_L2",
start_time="2024-07-01T10:00:00Z",
end_time="2024-07-01T14:00:00Z",
grid={"site": "Granada", "size_yx": [100, 100], "resolution_m": 10000.0},
)
Tip
examples/all_options.py lists every parameter each gatherer accepts, with its default and valid values.
Pipeline¶
The Pipeline class orchestrates multiple tasks:
from gathering import Task, Pipeline
pipeline = Pipeline([t1, t2, t3, t4])
# Preview in Jupyter (renders rich HTML table)
pipeline
Sequential Execution¶
Parallel Execution & Concurrency Levels¶
Gathering provides two distinct levels of concurrency that can be combined:
-
Task-Level Concurrency (Pipeline): Run multiple independent source gatherers in parallel threads:
-
Granule-Level Concurrency (Gatherer): Download multiple granules or product chunks concurrently within a single source task:
Accessing Results¶
# List of downloaded file paths
print(results.files)
# Convert to pandas DataFrame
df_summary = results.to_dataframe()
Single Task Execution¶
You can also run a single task directly:
from gathering import Task
task = Task.aeronet(
station="Granada",
product="AOD15",
start_date="2023-07-01",
end_date="2023-07-05",
)
result = task.run()
print(result.files)
YAML Execution from Python¶
from gathering import run_from_yaml
# Execute a YAML configuration file
results = run_from_yaml("examples/multi_source_campaign_pipeline.yaml")
Shared Vocabulary¶
Ground-based sources (AERONET, Pandonia, ACTRIS Cloudnet, ACTRIS Ares) use unified parameter names. Deprecated aliases are accepted with a warning:
| Concept | Canonical Name | Deprecated Aliases |
|---|---|---|
| Measurement site | station |
site, site_id, stations |
| Product / data type | product |
data_type, fcodes, kind |
| Range start | start_date |
date_from, from_date |
| Range end | end_date |
date_to, to_date |
Deprecation Policy
Deprecated aliases are supported throughout the 0.x series and scheduled for removal in 1.0.