Metadata-Version: 2.4
Name: dafab_client
Version: 4.0.0
Summary: Python client for the DaFab catalogue of Sentinel-2 imagery and AI-derived water and agriculture products, built on Rucio.
Author: DaFab
License-Expression: Apache-2.0
Project-URL: Homepage, https://www.dafab-ai.eu
Project-URL: Discovery site, https://dafab.cern.ch/
Project-URL: STAC catalogue, https://dafab.cern.ch/stac
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Scientific/Engineering :: GIS
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: folium<1,>=0.20.0
Requires-Dist: jsonschema<5,>=4.26.0
Requires-Dist: packaging<27,>=26.1
Requires-Dist: requests<3,>=2.33.1
Requires-Dist: typing_extensions<5,>=4.15
Requires-Dist: urllib3<3,>=2.6.3
Provides-Extra: notebooks
Requires-Dist: ipykernel<7,>=6.30; extra == "notebooks"
Requires-Dist: jupyter<2,>=1.1; extra == "notebooks"
Provides-Extra: previews
Requires-Dist: Pillow<13,>=12.2.0; extra == "previews"
Requires-Dist: rasterio<2,>=1.4.4; extra == "previews"
Dynamic: license-file

# DaFab Client

The Python client for the DaFab catalogue.

[DaFab](https://www.dafab-ai.eu) processes Copernicus Sentinel-2 imagery into AI-generated water-analysis and field-delineation products on European HPC systems, and publishes the images and the products as STAC items in a Rucio catalogue hosted at CERN. This client reads that catalogue, searches it with typed metadata filters, follows the lineage between a product and the image it was derived from, and downloads assets. It works out of the box with a shared read profile, no configuration needed.

You can also browse the same catalogue without code, on the [discovery site](https://dafab.cern.ch/) or through its [STAC root](https://dafab.cern.ch/stac).

## Install

```bash
pip install dafab-client
```

Python 3.10 or newer. The client talks to `dafab.cern.ch` over HTTPS and downloads assets from the project's object store. For the notebook, install the `notebooks` extra described below.

## First call

```python
import dafab_client as dc

dc.ping()    # the server version
dc.whoami()  # the shared read profile, account user_dafab
```

The package ships the profile `user_dafab`, meant for reading the public catalogue, together with the TLS trust bundle it uses. Publication needs a personal account, see Publishing below.

## What the catalogue holds

DaFab's images and products live in the scope `dafab` as STAC 1.1 items, each stored as a Rucio container that carries its full JSON document.

| Collection | Content | Example item |
|---|---|---|
| `sentinel_2_l2a` | Sentinel-2 L2A source images and their previews | `S2A_MSIL2A_20160229T032652_N0500_R018_T48QUH_20231011T191655` |
| `water_analysis` | water maps, excess and deficit polygons, previews | the image id followed by `_water_analysis_300` |
| `smart_agriculture` | field boundaries and previews | the image id followed by `_smart_agriculture_300` |

A product carries a `derived_from` link to its source image, and the catalogue attaches each product below its source image and its facet catalogue, so lineage can be followed both ways. Facet catalogues such as `agriculture_season`, `water_anomaly` and `water_basin` group products by season, by flood, drought or normal conditions, and by river basin.

Reference collections from other providers live in their own scopes beside `dafab`, one collection each. Their documents keep the metadata and the links of their providers.

| Scope | Collection | Content |
|---|---|---|
| `hand` | `glo-30-hand` | Height Above Nearest Drainage, 1 degree tiles at 30 m |
| `worldcover` | `esa-worldcover-map-10m-2021-v2` | ESA WorldCover 2021 land cover, 3 degree tiles at 10 m |
| `gfm` | `GFM` | reference water masks of the Copernicus Global Flood Monitoring |
| `fiboa` | `fiboa` | open field boundary datasets published by fiboa |

The helpers read the scope `dafab` unless you pass `scope=`. The STAC root covers `dafab`, and the discovery site shows all five scopes.

```python
dc.list_stac_scopes()                       # prints the scopes visible to the profile
dc.get_catalogs_and_collections()           # collections and facet catalogues of dafab, as {"scope", "id"} rows
dc.get_catalogs_and_collections(scope="hand")                    # [{"scope": "hand", "id": "glo-30-hand"}]
dc.get_items(filters={"name": "S2A_MSIL2A_20160229T032652_N0500_R018_T48QUH*"}, long=False)   # one tile and date, the image and its two products
dc.get_items_by_bbox(-17.5, 15.2, -17.4, 15.3, scope="hand")     # the HAND tile under a small box
```

`get_items` returns every item of the scope without a filter, which in `dafab` is more than 17,000 ids and takes about 20 seconds, since the server answers in pages of 100. Narrow it with `name` patterns (`*` or `%` match any sequence) or use the searches below.

## Search

Searches are JSON filters made of comparison nodes, each with a `path` into the document, a `comparator`, a `value` and a `valueType`, combined by logical nodes with `and`, `or` and `not`. The server evaluates any path of the stored documents, and the paths that the catalogue declares, such as `collection`, `properties.datetime` and `properties.eo:cloud_cover`, have indexes and answer fastest.

```python
clear_images_2016 = {
    "type": "logical", "combinator": "and", "filters": [
        {"type": "comparison", "comparator": "equals", "path": ["collection"], "value": "sentinel_2_l2a", "valueType": "string"},
        {"type": "comparison", "comparator": "lessThanOrEqual", "path": ["properties", "eo:cloud_cover"], "value": 20, "valueType": "number"},
        {"type": "comparison", "comparator": "greaterThanOrEqual", "path": ["properties", "datetime"], "value": "2016-01-01T00:00:00Z", "valueType": "date"},
        {"type": "comparison", "comparator": "lessThan", "path": ["properties", "datetime"], "value": "2017-01-01T00:00:00Z", "valueType": "date"},
    ],
}
ids = dc.get_items_by_enhanced_filter(clear_images_2016, return_mode="ids")
items = dc.get_items_by_enhanced_filter(clear_images_2016, return_mode="metadata")   # full documents
```

A group marked `"inherited": true` is evaluated against the ancestors of each candidate, so a search can select products by properties of their source image. This finds field-delineation products whose source image had at most 20 percent cloud cover.

```python
fields_from_clear_images = {
    "type": "logical", "combinator": "and", "filters": [
        {"type": "comparison", "comparator": "equals", "path": ["collection"], "value": "smart_agriculture", "valueType": "string"},
        {"type": "logical", "combinator": "and", "inherited": True, "filters": [
            {"type": "comparison", "comparator": "equals", "path": ["collection"], "value": "sentinel_2_l2a", "valueType": "string"},
            {"type": "comparison", "comparator": "lessThanOrEqual", "path": ["properties", "eo:cloud_cover"], "value": 20, "valueType": "number"},
        ]},
    ],
}
product_ids = dc.get_items_by_enhanced_filter(fields_from_clear_images)
```

| Comparators | Meaning |
|---|---|
| `equals`, `notEquals` | equality, or inequality |
| `lessThan`, `lessThanOrEqual`, `greaterThan`, `greaterThanOrEqual` | ordered comparison of numbers, dates or strings |
| `in` | the value is one of a list |
| `like`, `notLike` | text pattern, `*` or `%` for any sequence and `?` for one character |
| `contains`, `containsAny`, `containsAll` | array membership |
| `isNull`, `isNotNull`, `isEmpty`, `isNotEmpty` | null and empty-array checks, written without a `value` member |

Value types are `string`, `number`, `integer`, `boolean`, `date` and `array`. A document that lacks the field matches neither a typed comparison nor its negation, nor `isNull` or `isNotNull`, so incomplete metadata is left out rather than silently included.

Three helpers build the common spatial and temporal searches for you, and return ids or full documents with `return_mode`. A time range matches an item whose `properties.datetime` falls in it, or whose `start_datetime` to `end_datetime` span overlaps it, such as a HAND tile, as the discovery site does.

```python
dc.get_items_by_timerange("2016-02-01T00:00:00Z", "2016-03-01T00:00:00Z")
dc.get_items_by_bbox(103.5, 19.8, 104.2, 20.8)                    # west, south, east, north
dc.get_items_by_bbox_and_timerange([103.5, 19.8, 104.2, 20.8], ["2016-02-01T00:00:00Z", "2016-03-01T00:00:00Z"])
```

Lineage and facets have their own helpers.

```python
image = "S2A_MSIL2A_20160229T032652_N0500_R018_T48QUH_20231011T191655"
product = image + "_smart_agriculture_300"

dc.get_source_original_item_ids_from_derived_item(product)      # the image the product came from
dc.get_related_item_ids_from_original_item(image)                # every product of the image
dc.get_sibling_derived_item_ids(image, collection_id="smart_agriculture")
dc.get_item_ids_by_top_facet_catalog("water_basin")              # products placed under any river basin
dc.get_item_ids_by_facet_value_catalog("water_basin_mekong")     # products in one basin
dc.get_item_facet_placements(product, collection_id="smart_agriculture")
```

Project members find the full reference of comparators, types, null handling, sorting and paging in the repository's `docs/enhanced-filtering.md`.

## Read metadata

```python
document = dc.get_bulk_metadata(product)                          # the whole STAC item
subset = dc.get_bulk_metadata(product, paths=["/collection", "/properties/datetime", "/assets"])
dc.extract_metadata_value("/properties/datetime", pname=image)     # one value by JSON pointer, as text
report = dc.list_item_asset_entries(product, check_storage=True)  # every asset and whether its file is in storage
```

`dc.as_json(document)` formats any of these for display.

## Download assets

```python
dc.download_item_asset(product, "dafab-field-boundaries-thumbnail", destination_dir="downloads")
href = dc.build_stable_asset_href(image, "TCI_20m")               # the permanent URL of an asset
dc.download_asset_from_stable_href(href, destination_dir="downloads")
dc.download_all_derived_item_assets(product, destination_dir="downloads", overwrite=True)   # the thumbnail above again, with the other assets
```

Downloads go to `destination_dir`, the current directory by default, and refuse to replace an existing file unless `overwrite=True`. Where a local copy of the storage is mounted, `download_item_asset` uses it and reports `mode == "local_posix"`.

`dc.get_map(items, output_path="map.html")` renders the footprints of a list of documents to an interactive HTML map.

## Notebook

The packaged notebook walks through most of the above against the live catalogue, the reference collections included, and runs unchanged with the shared profile. It takes about three minutes, and its download cells write about 175 MB under `demo-data/` in the current directory.

```bash
pip install "dafab-client[notebooks]"
python -c "import dafab_client as dc; print(dc.get_example('user'))"
jupyter notebook dafab_simple_user_workflows.ipynb
```

## Validate a document

The validators check an item against the schemas shipped with the client, either a local file or the stored copy on the server, without changing anything.

```python
dc.validate_original_item(source="server", expected_item_id=image, expected_collection_id="sentinel_2_l2a")
dc.validate_derived_item(source="local", item_path="my_product.json")   # a product document of your own
```

## Configuration

Nothing is required for reading. These variables change the defaults.

| Variable | Effect |
|---|---|
| `DAFAB_PROFILE` | the account whose profile to load, default `user_dafab` |
| `DAFAB_PROFILE_PATH` | an explicit profile file |
| `DAFAB_PROFILE_DIR` | the directory holding `<account>/config`, default `~/.config/dafab/credentials/profiles` or `$XDG_CONFIG_HOME/dafab/credentials/profiles` |
| `DAFAB_DEMO_DATA_DIR` | the folder under which `get_map` writes `filters/bbox_map.html` when it gets no `output_path`, default `./demo-data` |

Set them before importing the client. A profile is a JSON file with the connection and the credentials of one account.

```json
{
  "rucio_host": "https://dafab.cern.ch",
  "auth_host": "https://dafab.cern.ch",
  "auth_type": "userpass",
  "account": "my_account",
  "creds": {"username": "my_account", "password": "..."},
  "ca_cert": "$client_path/_rucio/secrets/user_dafab/rucio_ca_bundle.pem",
  "vo": "def"
}
```

`ca_cert` names the TLS trust bundle. `$client_path` is the installed package, `$client_account` the account name, and `$profile_path` the directory of the profile file, so a personal profile can reuse the shipped bundle as above. Keep profile files private.

## Publishing

Publishing images and products, attaching them to their collections and facets, and repairing collection extents are done by DaFab's processing pipelines with the publication helpers of this package, `ensure_item`, `set_metadata`, `publish_original_asset`, `publish_derived_asset` and the `sync_*_collection_extent` family. They write to the catalogue and need an account with write access. `prepare_original_item_metadata`, which makes the thumbnail and overview previews of an image, needs the `previews` extra, `pip install "dafab-client[previews]"`. Project members find the procedures in the repository's `docs/`. To request an account, contact the project through [dafab-ai.eu](https://www.dafab-ai.eu).

## Licence

Apache License 2.0. The package bundles the Rucio client modules, which are also Apache 2.0.
