Metadata-Version: 2.4
Name: rapidsar
Version: 0.0.1a1
Summary: Rapid Sentinel-1 SAR building damage-change detection and mapping
Author: RapidSAR contributors
License-Expression: MIT
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Operating System :: OS Independent
Classifier: Development Status :: 3 - Alpha
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: affine>=2.4
Requires-Dist: numpy>=1.24
Requires-Dist: scipy>=1.10
Requires-Dist: rasterio>=1.3
Requires-Dist: shapely>=2.0
Requires-Dist: pyproj>=3.5
Requires-Dist: matplotlib>=3.7
Requires-Dist: pyshp>=2.3.1
Dynamic: license-file

# rapidsar 0.0.1a1

RapidSAR is a Python package for **rapid building damage-change detection
from Sentinel-1 SAR**. Its original `run_pipeline` workflow takes two or three
pre-event SLC scenes and one post-event SLC scene, then produces building-level
change scores, binary change-screen flags, a static map, and a self-contained
interactive map. The design keeps the input and feature interface stable so
other detection models can be added and compared.

The package currently offers two label-free methods:

- `two_stage` (default): use pre-event pseudo-events to mark high-change and
  stable examples, then train Stage 2 on those pseudo-labels using canonical
  temporal, SAR and terrain features. Spatial cross-fitting is preferred;
  when spatial folds are too small, RapidSAR retries without the spatial
  buffer and then uses a full-data fit. Duplicate aliases and the direct Stage
  1 score are excluded from Stage 2.
- `unsupervised`: use the same pre-event calibration and building features,
  but stop after Stage 1. This is a complete method and does not fit Stage 2.

Both methods screen for SAR change. `binary_change` means the building passed
the selected change screen; it does not mean confirmed destroyed. The score is
not a calibrated probability of damage, so `destruction_status` remains
`undetermined` and `destruction_confidence` remains null.

## Method strictness and Stage 2 support

`PipelineConfig` exposes the method's strictness controls directly. A higher
`positive_quantile`, lower `negative_quantile`, higher `min_change_db`, and
lower `max_stable_change_db` select more conservative pseudo-labels. For the
final Stage 2 candidate flag, increase `stage2_score_threshold` above its
default of 0.5. These settings change screening behavior; scores still are
not calibrated probabilities of damage.

Stage 2 has an absolute minimum of 10 high-change and 10 stable pseudo-labels.
The preferred support is 50 per class. With the defaults, counts from 10
through 49 per class still run and emit a Python runtime warning; the warning
is also saved in `analysis_summary.json` and the GeoJSON collection's
`run_warnings`. Below the configured minimum, RapidSAR reports the counts and
returns the Stage 1 result.
If spatial folds do not meet the minimum, it first retries with a 0 m spatial
buffer, then with a full-data fit. The latter is reported because its scores
for pseudo-labeled buildings are in-sample and can be optimistic.

## Install and verified environment

Package metadata permits Python 3.10 or newer and uses minimum supported
versions for its Python dependencies. These ranges are intended to support
Windows, macOS, and Linux. The reference environment below was tested for
processed-raster scoring, maps, evaluation, and footprint retrieval. The
September 2026 review also checked SNAP graph generation and installed operator
parameters; it did not repeat a full raw-SLC processing run. Other platforms
and the minimum dependency versions have not yet been tested. For the exact
Windows Python environment, use the included platform-specific lock file.

| Component | Exact reference-tested version |
| --- | --- |
| Python | CPython 3.12.14, 64-bit Windows (AMD64) |
| NumPy | 2.5.3 |
| SciPy | 1.18.1 |
| Rasterio | 1.5.1 |
| Shapely | 2.1.2 |
| PyProj | 3.8.0 |
| Matplotlib | 3.11.2 |
| Affine | 3.0.1 |
| setuptools (build) | 84.0.0 |
| wheel (build) | 0.48.0 |
| ESA SNAP GPT | 14.0.0, external tool for raw SLC processing |
| Java | SNAP-bundled Java 21.0.6 |

For the first PyPI prerelease, install with:

```bash
python -m pip install --pre rapidsar==0.0.1a1
```

Or install from a source checkout with:

```bash
python -m pip install .
```

To reproduce the verified Windows environment with `uv`:

```powershell
uv venv --python 3.12.14 .venv
uv pip sync --python .venv/Scripts/python.exe requirements-rapidsar-py312-win-amd64.lock
uv pip install --python .venv/Scripts/python.exe --no-deps -e .
```

The lock captures the exact Windows reference environment; it is not the
package's cross-platform dependency specification. SNAP is not installed by
pip. It is needed only when the package must process SLC SAFE ZIPs. SNAP's GPT
executable and its DEM/orbit data access must be available on the machine.
The tested SNAP installation bundled Java, so a separate Java installation
was not needed. If usable processed SNAP GeoTIFFs are already available,
scoring and map generation do not call SNAP.

## Main workflow

Pass two or three pre-event acquisitions, one post-event acquisition, and a
WGS84 footprint GeoJSON:

```python
from rapidsar import PipelineConfig, run_pipeline

result = run_pipeline(
    pre_fire_scenes=["pre_1.SAFE.zip", "pre_2.SAFE.zip", "pre_3.SAFE.zip"],
    post_fire_scene="post.SAFE.zip",
    footprints="buildings.geojson",
    output_dir="outputs/event",
    method="two_stage",  # or "unsupervised"
    config=PipelineConfig(snap_gpt="/path/to/snap/bin/gpt"),
)
print(result.to_dict())
```

If you do not have building footprints, pass an AOI instead. The default
`footprint_source="microsoft"` downloads Microsoft's legacy US state archive
or global country tiles, selects whole footprints that intersect the AOI, then
scores each building. For a US AOI, supply `microsoft_state` if the AOI lacks a
`state`, `state_name`, `state_code`, `STUSPS`, `postal`, or `admin1` property.
Elsewhere, supply `microsoft_country` if it lacks a `country` or `country_name`
property:

```python
result = run_pipeline(
    pre_fire_scenes=["pre_1.SAFE.zip", "pre_2.SAFE.zip", "pre_3.SAFE.zip"],
    post_fire_scene="post.SAFE.zip",
    aoi="study_area.geojson",  # WGS84; a Polygon or FeatureCollection
    microsoft_state="California",  # e.g. "Hawaii" for Maui, or a USPS code
    # output_dir defaults to ./rapidsar_output
)
```

For a global AOI such as Turkey, use `microsoft_country="Turkey"` instead.
Country names ignore spacing and punctuation, and `Türkiye` is accepted.
An explicit country selects the global source, including `"United States"`;
US state metadata selects the legacy source when no country override is given.
The default global source is pinned to release `2026-08-13`; set
`microsoft_global_release` to a different dated index when that snapshot is
available from Microsoft. The CLI accepts the same values as
`--microsoft-country` and `--microsoft-global-release`.

The matching Microsoft state archive or global country tiles and index are
cached at `<output_dir>/inputs/source_archives/`; the AOI-filtered input layer and its
provenance record are written to
`<output_dir>/inputs/building_footprints_microsoft.geojson` and
`building_footprints_microsoft.provenance.json`. The scored building GeoJSON,
score GeoTIFF, maps, and summary remain at the output-folder root. Processed
rasters, SNAP graphs, and logs remain in `processed/`, `graphs/`, and `logs/`.
Downloaded US state ZIPs, global tile archives, and the global tile index are
retained for reuse. A California state archive is large; global downloads are
limited to country tiles intersecting the AOI.
On repeat runs, matching cached footprints are reused; an existing generated
result still requires `overwrite=True` before the pipeline replaces it.

Every default-sourced building has a `footprint_data_year` property. For US v2
it comes from Microsoft's per-building `capture_dates_range` when available;
otherwise the value is `null` and the basis field says the year was not
provided. For global footprints, Microsoft does not publish a capture year per
building, so that value is `null`; the metadata instead lists the pinned
dataset release and the collection-wide imagery range (`2014-2025` for the
verified 2026-08-13 release, varying by place). Other release dates without
verified imagery metadata explicitly report an unknown range. Year counts
and source metadata are copied into the scored GeoJSON,
analysis summary, and map annotations. The current global release includes
some 2025 imagery, so check its vintage against the event before using it as a
pre-event inventory. Microsoft also documents variable local coverage and
quality. See the [US source](https://github.com/microsoft/USBuildingFootprints)
and [global source](https://github.com/microsoft/GlobalMLBuildingFootprints)
for imagery vintage and license details. For raster-only analysis, use
`footprint_source="none"` with an AOI.

The run writes processed scene rasters, a scored `building_change.geojson`, a
change-score GeoTIFF, a Matplotlib PNG and SVG, a standalone HTML map, and JSON
summaries. GeoJSON preserves source footprint properties and adds `change_score`,
`change_score_kind`, `binary_change`, `change_status`, Stage 1 pseudo-label
fields, data-support fields, acquisition dates, and the explicitly undetermined
destruction fields.

## Use the stages independently

Raw SLC processing can be separated from scoring and mapping:

```python
from rapidsar import process_slc_archives, score_buildings
from rapidsar import (make_static_map, make_interactive_map,
                      make_interactive_map_with_basemap)

processed = process_slc_archives(
    pre_fire, post_fire, "outputs/event", footprints="buildings.geojson"
)
scored = score_buildings(
    [*processed.pre_fire, processed.post_fire], "buildings.geojson",
    "outputs/event/building_change.geojson",
    coherence_path=processed.coherence,
    method="unsupervised",
)
make_static_map(scored.geojson_path)
make_interactive_map(scored.geojson_path)
make_interactive_map_with_basemap(
    scored.geojson_path,
    reference_data_dir="outputs/event/map_reference_cache",
)
```

The reference-map function caches small Census TIGER/Line boundary files and
roads for counties touching the map. It clips and simplifies roads around the
building view before embedding them; the resulting HTML makes no tile or API
requests and works offline. The cache avoids repeat downloads.

If SNAP-processed GeoTIFFs already exist, use `score_buildings` (with
footprints) or `score_raster` (without footprints) directly. The expected
scene rasters are calibrated sigma0 with VV/VH, elevation, local incidence,
and layover bands. The input order is two or three pre-event scenes followed
by one post-event scene.

`run_pipeline`, `process_slc_archives`, scoring functions, and map functions
accept `overwrite=False` by default. Processed rasters are reused only when
their `<raster>.processing.json` receipt matches the processing graph, inputs
(resolved path, size, modification time), DEM, and output file identity.
Changing the AOI, scene or processing settings invalidates reuse. A raster
from an older run without this receipt can still be passed directly to
`score_buildings` or `score_raster`; automatic processing will require a new
output folder or `overwrite=True`. Completed SNAP products survive interrupted
runs and failed replacements. These identity checks are not full-file checksums.
Existing output maps, score files, and GeoJSON are protected from
replacement. Set `overwrite=True` to recompute cached SNAP products and replace
outputs. In the CLI, use `--overwrite`.

The raster-only pipeline also writes `analysis_summary.json` and
`pipeline_outputs.json`. The `change_score.tif` remains the unsupervised pixel
screen for both building methods; Stage 2 building scores are stored in the
GeoJSON. Scoring multiple building methods into one directory would share
the raster and summary filenames, so use a separate output directory per run.

## Add a model

All building methods receive the same SAR-derived per-building feature rows,
projected coordinates, and `PipelineConfig`. Implement a `score(rows, xy,
config)` method returning `ModelScores`, then either pass the model instance
to `score_buildings(..., model=my_model)` or register a factory:

```python
from rapidsar import register_model, run_pipeline

register_model("my_model", MyModel)
result = run_pipeline(..., method="my_model")
```

The input contract does not require a particular learning method. Built-in
choices are `TwoStageModel` and `UnsupervisedModel`; future methods can use
the same two-to-three pre-event and one post-event scene setup. Damage labels
are not passed to this model interface.
Use `NaN` for unavailable scores and `-1` for unknown binary flags in a custom
`ModelScores` result; unknown flags are written as GeoJSON `null`.

## Independent evaluation

When independent reference labels are available, evaluate a frozen scored
GeoJSON. The evaluator plots a ROC curve, returns AUC, precision, recall, F1,
accuracy, specificity, balanced accuracy, and the confusion matrix, and can
write the plot and metrics JSON. Labels are read only by this evaluation step;
they do not alter model scores or select thresholds.

```python
from rapidsar import evaluate_geojson

metrics = evaluate_geojson(
    "scored_buildings_with_reference_labels.geojson",
    truth_field="reference_destroyed",
    score_field="change_score",
    prediction_field="binary_change",  # evaluate the frozen binary screen
    output_path="outputs/event/roc_curve.png",
)
print(metrics["auc"], metrics["f1"], metrics["precision"], metrics["recall"])
```

When Stage 2 fits successfully, its score is in [0, 1] and the built-in binary
flag uses `stage2_score_threshold` (0.5 by default). If spatial folds are too
small, the model retries with a zero buffer and then without a spatial split;
Stage 1 is used only when the global label minimum is not met or fitting fails.
`change_score_kind` and `stage2_model_status` identify the selected path.
For `unsupervised`, scores are not probabilities; use the existing frozen
`binary_change` field for precision/recall/F1, or supply a threshold chosen
without the held-out evaluation labels. The ROC/AUC uses continuous scores in
either case. Evaluation excludes features with missing labels or scores and
requires both reference classes. It reports excluded rows separately from
rows lacking a usable binary prediction. Scores must increase toward the
positive reference class. Numeric 0/1 or boolean labels are accepted; map
textual inspection categories to those labels before evaluation.

The command-line equivalent is:

```bash
rapidsar evaluate scored_buildings_with_reference_labels.geojson \
  --truth-field reference_destroyed --output outputs/event/roc_curve.png
```

## Input and processing notes

- Inputs must be Sentinel-1 IW SLC dual-VV/VH (`1SDV`) SAFE ZIPs with standard
  filenames and a readable manifest giving the same platform and relative
  orbit; pre-event scenes are sorted by acquisition date. GRD, EW, mixed
  platforms/tracks and more than three pre-event scenes are not supported by
  the current public workflow.
- Building footprints and AOIs use RFC 7946 Polygon/MultiPolygon GeoJSON
  coordinates (WGS84). Analysis grids must be projected in metres. A local
  UTM/UPS CRS is inferred unless `target_crs` is supplied. Split areas crossing
  the antimeridian into separate local analyses.
- SNAP selects the IW swath whose SAFE annotation geolocation grids overlap
  the AOI most consistently (`subswath="auto"`, the default). Specify IW1,
  IW2 or IW3 to force a swath. Processing uses 10 m pixels by default and the
  nearest pre/post pair for optional VV coherence. Set `include_coherence=False`
  if that pair is unsuitable or unavailable. The requested and selected swath
  are recorded in `processing_metadata.json` and the run summary.
- Coherence coregistration retains the selected subswath's bursts so ESD has
  adjacent-burst overlap, then crops the output to the AOI. This can require
  substantially more memory and DEM coverage than the final output area.
  Single-burst or non-overlapping pairs still need separate handling; errors
  are reported rather than silently changing the feature set.
- With an external `dem_path`, nodata is read from the raster or supplied as
  `dem_nodata_value`. Use `dem_apply_egm=True` for the SNAP geoid-to-ellipsoid
  adjustment, or `False` when the DEM already has ellipsoidal heights. The
  same DEM settings are passed to coregistration and terrain correction.
- When footprints define the area, SNAP uses their bounding extent buffered
  by 250 m. A supplied AOI takes precedence.
- Disk, memory, DEM availability, registration quality, and valid raster
  support limit coverage. Buildings without sufficient SAR support remain
  unknown instead of being labeled unchanged. The valid-support fraction
  includes requested footprint-buffer pixels outside the raster. Missing or
  empty footprint geometries are omitted and counted in `omitted_geometry_count`.
- This release makes change-screen outputs for rapid inspection. It does not
  claim that a Sentinel-1 change uniquely identifies fire damage; independent
  assessment remains necessary.
