Metadata-Version: 2.4
Name: dhis2-fetch
Version: 0.1.2
Summary: Read-only Python client for the DHIS2 Web API, returning pandas DataFrames
Author-email: Valérian Turbé <valerian.turbe@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/Valerian8/dhis2-fetch
Project-URL: Repository, https://github.com/Valerian8/dhis2-fetch
Project-URL: Issues, https://github.com/Valerian8/dhis2-fetch/issues
Keywords: dhis2,health,hmis,pandas,api,data
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Healthcare Industry
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Typing :: Typed
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: requests
Requires-Dist: pandas
Requires-Dist: tqdm
Provides-Extra: dev
Requires-Dist: pytest>=7; extra == "dev"
Requires-Dist: responses>=0.23; extra == "dev"
Dynamic: license-file

# dhis2-fetch

A **read-only** Python client for the DHIS2 Web API that returns tidy `pandas` DataFrames, built for analysts working in notebooks and scripts.

It authenticates with **your own DHIS2 account** and can only access data that account is already permitted to see in the DHIS2 web interface. It issues only `GET` requests and cannot create, modify, or delete items on a DHIS2 instance.

To help users struggling with intermittent internet connection and/or unstable DHIS2 instances, long downloads are resumable: progress is written to a local SQLite file as it goes, so an interrupted job continues where it stopped instead of starting over.

**Who this is for:** Bulk analysis of DHIS2 data is easiest against the database directly, but most analysts have no SQL access to it. The database sits behind the instance and making SQL queries is usually restricted to admins. This package uses the Web API with the account you already have and is designed to be practical for data analysis: it handles the batching, paging, retries and reshaping, and hands back a clean DataFrame you can work with, without needing access to database credentials.

It is also useful for **automated pipelines** that pull data or saved visualizations from DHIS2 on a schedule, for example to feed a dashboard or a database. Anything that would prompt (credentials, dataset choices, resuming a job) can be passed as an argument instead; interrupted runs resume, and failed requests are recorded in `df.attrs` rather than stopping the run, so a pipeline can check them and re-run to retry.

---

## Install

```bash
pip install dhis2-fetch
```

Requires Python 3.9+, and installs `requests`, `pandas` and `tqdm`.

## Quick start

Every example on this page runs as written against the public DHIS2 demo, so you can try the package
before pointing it at your own instance:

```python
import dhis2_fetch as dfetch

dfetch.connect('https://play.im.dhis2.org/stable-2-43-1', 'admin', 'district')   # the public demo
```

```
Connected to https://play.im.dhis2.org/stable-2-43-1 as admin
```

On your own instance, you can leave the credentials out and you will be prompted for them, so they don't sit in
your script:

```python
dfetch.connect('https://your-dhis2-instance.org')      # prompts for username and password
```

The demo is a shared server that is reset regularly, so its data moves and it is sometimes down. Its
analytics tables are also not always built; when they are missing, the aggregate data and visualization
examples below come back empty or failed while the events and maps examples still work.

Every download is asked for at an org unit level, and often for particular org units, so start with how
to find those. Then follow the recipe for what you need; each one goes in the same order: **find the
UID, download, look at the result**.

1. **Aggregate data** — indicators, data elements and datasets
2. **Events** — one row per event from a tracker or event program
3. **Saved visualizations** — the data behind a chart or pivot table
4. **Boundaries and maps** — GeoJSON per org unit level, joined to your data

## Org units: levels and UIDs

The levels tell you what to pass as `download_level`:

```python
dfetch.fetch_orgunit_levels()
# {1: 'National', 2: 'District', 3: 'Chiefdom', 4: 'Facility'}
```

`fetch_orgunit_tree(level)` lists every org unit at a level, with the name and UID of each level above
it. Search it by name to find the UIDs to pass as `filter_ou_uids`:

```python
chiefdoms = dfetch.fetch_orgunit_tree(3)
chiefdoms[chiefdoms['Chiefdom'].str.contains('badj', case=False)]
#        National National_uid District District_uid Chiefdom Chiefdom_uid
# 0  Sierra Leone  ImspTQPwCqd       Bo  O6uvpzGd5pu   Badjia  YuQRtpLP10I
```

Add `boundary_ou_uid=` to list only what sits beneath one org unit, for example the facilities in that
chiefdom:

```python
dfetch.fetch_orgunit_tree(4, boundary_ou_uid='YuQRtpLP10I')[['Chiefdom', 'Facility', 'Facility_uid']]
#   Chiefdom       Facility Facility_uid
# 0   Badjia   Ngelehun CHC  DiszpKrYNg8
# 1   Badjia  Njandama MCHP  g8upMTyEZGZ
```

The columns are the same hierarchy columns every download carries, so the tree also works as a lookup
table: `dfetch.fetch_orgunit_tree(4)` is every facility on the instance with its
chiefdom and district, ready to merge onto your own data by `Facility_uid`.

## 1. Aggregate data

Find the UID of what you want by searching the instance's data items by name:

```python
items = dfetch.fetch_data_items()
items[items['dataname'].str.contains('ANC 1', case=False)]
#      datatype       dataid                dataname
# 20         DE  fbfJHSPpUQD           ANC 1st visit
# 350        DE  ybzlGLjWwnK  LLITN given at ANC 1st
# 1573        I  ReUHfIn0pTQ    ANC 1-3 Dropout Rate
# 1574        I  Uvn6LCg7dVU          ANC 1 Coverage
```

`datatype` is `DE` (data element), `I` (indicator), `PI` (program indicator) or `DS` (dataset). Pass
`verbose=True` to also get `valueType`/`aggregationType` for data elements and `numerator`/`denominator`
for indicators.

Then download it, at one of the levels from `fetch_orgunit_levels()` above:

```python
df = dfetch.download_data(
    uids=['Uvn6LCg7dVU'],          # ANC 1 Coverage, from the search above
    start_period='202501',
    end_period='202512',
    period_unit='Month',
    download_level=3,              # Chiefdom
)

df.head(3)
#    Year  Month      pe  ...  District District_uid Chiefdom Chiefdom_uid  ANC 1 Coverage
# 0  2025      1  202501  ...        Bo  O6uvpzGd5pu   Badjia  YuQRtpLP10I          140.43
# 1  2025      1  202501  ...        Bo  O6uvpzGd5pu    Baoma  vWbkYPRmKyS          179.75
# 2  2025      1  202501  ...        Bo  O6uvpzGd5pu   Bargbe  dGheVylzol6          138.63

df.to_csv('anc1_coverage_2025.csv', index=False)
```

`df` is a wide DataFrame: one row per period × org unit (here 158 chiefdoms × 12 months), one column per
data item, with the full org unit hierarchy joined on as a name and a `_uid` column for each level above
it. It also carries the org unit's group sets as columns (`Location Rural/Urban` on the demo); see "Org unit
attribute columns" below.

Pass several UIDs to get several columns. Datasets, other period units, resuming and naming jobs are
covered in the reference below.

## 2. Events

Find the program:

```python
programs = dfetch.fetch_programs()      # one row per program × stage × data element
programs[programs['program_name'].str.contains('Malaria case registration')][
    ['program_uid', 'program_name', 'stage_uid', 'stage_name']].drop_duplicates()
#      program_uid               program_name    stage_uid                 stage_name
# 108  VBqh0ynB2wv  Malaria case registration  pTo4uMt3xur  Malaria case registration
```

Events are requested one facility at a time, so pick an area: here Badjia chiefdom and its two
facilities, found in "Org units: levels and UIDs" above. Then download:

```python
events = dfetch.download_event_data(
    program_uid='VBqh0ynB2wv',          # Malaria case registration, from the search above
    start_date='2025-01-01',
    end_date='2025-03-31',
    download_level=4,                   # Facility: the level events are captured at
    filter_ou_uids=['YuQRtpLP10I'],     # Badjia chiefdom, from fetch_orgunit_tree above
)

events.shape
# (45, 18)
events[['event_date', 'Facility', 'Gender', 'Age in years']].head(3)
#   event_date      Facility  Gender  Age in years
# 0 2025-01-01  Ngelehun CHC  Female            14
# 1 2025-01-03  Ngelehun CHC    Male            63
# 2 2025-01-04  Ngelehun CHC  Female             4
```

One row per event, with the org unit hierarchy and attribute columns joined on, the event date split into
`Year`/`Month`, and each data element as a named column. Pass `stage_uid=` to restrict to one stage.
Resumable in the same way as aggregate downloads, with the same `df.attrs` (`events_available`,
`db_path`, `job_name`, `created_date`).

Two things differ from aggregate downloads:

- **`download_level` is the level events were captured at**, usually the facility level. Events are read
  from those exact org units, not from their descendants, so asking at district level returns only events
  registered against the districts themselves.
- **It makes one request per org unit**, so a whole country's facilities means thousands of requests. Pass
  `filter_ou_uids` to scope it to the units you need, as above; you will get a warning above 200.

## 3. Saved visualizations

Any chart or pivot table saved on the instance can be downloaded as data. Find it by name:

```python
viz = dfetch.fetch_visualizations()
viz[viz['name'].str.contains('ANC', case=False)][['id', 'name', 'type']].head(3)
#             id                                              name            type
# 0  LW0O27b7TdD                      ANC: 1-3 dropout rate Yearly          COLUMN
# 1  DkPKc1EUmC2               ANC: 1-3 trend lines last 12 months            LINE
# 2  IvXcdp2cFHa  ANC: 1-4 visits by districts this year (stacked)  STACKED_COLUMN

chart = dfetch.download_visualization('LW0O27b7TdD')     # ANC: 1-3 dropout rate Yearly
chart.head(3)
#                    Data Organisation unit  Value
# 0  ANC 1-3 Dropout Rate                Bo  34.48
# 1  ANC 1-3 Dropout Rate           Bombali  38.01
# 2  ANC 1-3 Dropout Rate            Bonthe  33.63
```

A chart comes back as one row per value; a pivot table comes back laid out as it is in DHIS2. Saved
visualizations often use relative periods such as "last 12 months", which DHIS2 resolves on the day you
download, so the periods you get change over time.

## 4. Boundaries and maps

```python
shapes = dfetch.fetch_shapefiles(levels=[3])      # leave levels out for every level
shapes
# ShapefileSet(Chiefdom: 158 features)

dfetch.save_shapefiles(shapes, output_dir='./outputs')
# {'Chiefdom': './outputs/Chiefdom.geojson'}
```

The result is an ordinary dict: `shapes['Chiefdom']` is a GeoJSON `FeatureCollection` you can hand
straight to geopandas, folium or `json.dump`. It only prints as a summary, because the real output can be
hundreds of thousands of coordinates. `save_shapefiles` writes one `.geojson` per level, creates
`output_dir` if needed, makes level names safe to use as filenames, and returns the paths. Levels without geometry are
skipped with a warning.

Each feature's properties use the same column names as the downloads (`Chiefdom`, `Chiefdom_uid`), so a
map of the data from recipe 1 is a merge away. This needs geopandas and matplotlib, which are not
dependencies of this package (`pip install geopandas matplotlib`):

```python
import geopandas as gpd

gdf = gpd.GeoDataFrame.from_features(shapes['Chiefdom'])
jan = df[df['pe'] == '202501'][['Chiefdom_uid', 'ANC 1 Coverage']]
gdf.merge(jan, on='Chiefdom_uid').plot(column='ANC 1 Coverage', legend=True)
```

---

# Reference

## Useful functions

| Function | What it gives you |
|---|---|
| `connect(url, username=None, pwd=None, quiet=False)` | Authenticates once; everything after uses that session |
| `active()` | The connection `connect()` made — check which instance and account you are on |
| `fetch_data_items(verbose=False)` | Every dataElement / indicator / programIndicator / dataSet on the instance — this is how you find UIDs before downloading data |
| `fetch_orgunit_levels()` | `{1: 'National', 2: 'Regional', ...}` — tells you what `download_level` to ask for |
| `fetch_orgunit_tree(level, boundary_ou_uid=None)` | Every org unit at a level, with the name and UID of each level above it — this is how you find org unit UIDs |
| `download_data(...)` | The main event: UIDs + period range + level → a wide DataFrame |
| `fetch_shapefiles(boundary_ou_uid=None, levels=None)` | `{level_name: GeoJSON FeatureCollection}` for mapping |
| `save_shapefiles(shapes, output_dir='./outputs')` | Writes those out as one `.geojson` per level; returns the paths |
| `fetch_visualizations()` | List all saved visualizations on the instance |
| `validate_visualization_uid(uid, viz_list=None)` | Resolves a visualization UID to `(name, object_type)` |
| `download_visualization(uid, object_type)` | The data behind a saved visualization, as a DataFrame — a pivot table as laid out, a chart as one row per value |
| `fetch_programs()` | Tracker programs, their stages, and each stage's data elements (one row per data element) |
| `download_event_data(...)` | A program's events as a wide DataFrame, one row per event |
| `download_events(...)` | The lower-level primitive behind it: raw events into the job database |

Lower-level building blocks (`download_data_analytics`, `select_dataset_elements`, `consolidate_data`,
`pivot_data`, `generate_periods`, …) are available in their modules in case you need to drive a step yourself: `dhis2_fetch.analytics`,
`.orgunits`, `.periods`, `.visualizations`, `.events`.

## The connection

`connect()` with no arguments prompts for the URL too. You can pass credentials explicitly
(`dfetch.connect(url, username, password)`) for unattended runs; pass `quiet=True` as well to leave out
the "Connected to" line, e.g. in a scheduled script. A failed login is still reported. Calling anything
before connecting raises `NotConnectedError`.

To check later which connection is in use:

```python
dfetch.active().url          # 'https://your-dhis2-instance.org'
dfetch.active().username     # 'your username'
```

`active()` raises `NotConnectedError` if you have not connected yet. There is one connection per Python
process: calling `connect()` again replaces it rather than adding a second. `connect()` also returns the
connection, so `client = dfetch.connect(...)` is equivalent, and if you need two instances live at
once, build `Dhis2Client` objects directly instead.

## Org unit levels and filtering

`download_level=4` downloads at the fourth level of the hierarchy. Indicators and data elements are
added up by DHIS2 from the levels below, so asking at district level gives district totals.

A dataset's data elements are different: they are read only where their values were entered, usually
the facilities, and are not added up. Ask for a dataset at that level, or pass its data element UIDs
themselves to have them added up (with `disaggregation=True` to keep the category breakdown). You get
a warning when a dataset comes back empty for this reason.

To restrict to part of the hierarchy, pass `filter_ou_uids=['uid1', 'uid2']` — these may be at the
download level itself (downloads exactly those units) or above it (downloads everything beneath them).
That second form is how you pull a sub-hierarchy: passing hospital UIDs with `download_level` set to the
level below returns every department of those hospitals.

To find the UIDs to pass, use `fetch_orgunit_tree`, as in "Org units: levels and UIDs" above.

## Org unit attribute columns

Every result also carries the org unit's **group set** memberships as columns, one per group set defined on the instance, describing the units at `download_level`. Pass `ou_attributes=False` to omit them.

These are the columns that tell you *what kind of thing* each row is, which matters because a level often mixes different kinds of unit. For example: on one instance, level N could hold 50 health districts **and** 700+ hospital departments, because hospitals are registered a level up, beside provinces, rather than inside districts. The group sets separate them cleanly:

| Column | Health districts | Hospital departments |
|---|---|---|
| `District_attribute` | `HD` | blank |
| `Type` | blank | `Hospital department` |

```python
df = dfetch.download_data(uids=[...], download_level=4, ...)

geographic = df[df['District_attribute'] == 'HD']        # the 50 districts
hospital   = df[df['Type'] == 'Hospital department']          # the 700+ hospital departments
```

Two things worth knowing before you filter:

- **The two sets are disjoint and both real.** Check whether they sum exactly to the national total: if
  they do, filtering to one genuinely drops the other's data rather than removing duplication. Hospital
  departments can be a small share of a routine indicator but a large share of specialist activity.
- **Mixing is not universal.** On the same instance, levels 5 and 6 are entirely geographic — the
  hospital branch has no children — so nothing needs filtering there. Check your own instance before
  assuming either way.

Group set names and values are defined per instance, so yours will differ; `District` and `Type` above are
that instance's own. The column is suffixed `_attribute` only when a group set shares its name with an org
unit level, as `District` does here.

## Periods

`period_unit` must be one of:

| Unit | Example | Unit | Example |
|---|---|---|---|
| `Year` | `2024` | `Weekly` | `2024W1` |
| `Month` | `202401` | `BiWeekly` | `2024BiW1` |
| `Quarter` | `2024Q1` | `Daily` | `20240101` |
| `BiMonth` | `2024B1` | `FinancialApril` | `2024April` |
| `SixMonth` | `2024S1` | `FinancialJuly` | `2024July` |
| `SixMonthApril` | `2024AprilS1` | `FinancialOct` / `FinancialNov` | `2024Oct` |

`start_period` and `end_period` use the same format, and every period in between is generated for you.
Periods with no data are still present as rows, with `NA` values.

## Datasets and reporting metrics

If any UID is a **dataSet**, you are asked what to pull from it — its data elements, its reporting
metrics, or both. There are five reporting metrics: `REPORTING_RATE`, `REPORTING_RATE_ON_TIME`,
`ACTUAL_REPORTS`, `ACTUAL_REPORTS_ON_TIME` and `EXPECTED_REPORTS`. To skip the prompt and run
unattended, answer in advance:

```python
df = dfetch.download_data(
    uids=['Uvn6LCg7dVU', 'vc6nF5yZsPR'],
    start_period='202401', end_period='202412', period_unit='Month',
    download_level=4,
    dataset_choices={'vc6nF5yZsPR': {'elements': 'all', 'metrics': True}},
)
```

`elements` accepts `'all'`, an explicit list of data element UIDs, or `None` for metrics only. Reporting
metric columns are placed after the raw data columns.

## Resuming an interrupted download

Each job writes to `./outputs/<job_name>.db`. Re-running the same call picks up where it left off and
tells you how many requests were already done. You will be asked to confirm; pass `force=True` to skip
that prompt.

If the job's defining parameters changed (instance URL, org units, period unit, disaggregation), it
stops rather than mixing incompatible data — delete the `.db` file to start fresh. The database also
means you can inspect exactly what was retrieved, and it is safe to delete once you have your DataFrame.

Useful extras on the result:

```python
df.attrs['data_available']   # {'Yes': [...], 'No': [...], 'Failed': [...], 'Skipped': [...]}
df.attrs['db_path']          # where the job database lives
df.attrs['job_name']
df.attrs['created_date']
```

### Naming a job yourself

The job name is built from the country, org units, level and period unit, so two downloads that differ
only in their UIDs or period range want the same file. The second one stops rather than mixing data.
Pass `job_name` to keep them apart:

```python
df_2023 = dfetch.download_data(uids=[...], start_period='202301', end_period='202312',
                            job_name='malaria_2023', ...)
df_2024 = dfetch.download_data(uids=[...], start_period='202401', end_period='202412',
                            job_name='malaria_2024', ...)
```

Each gets its own `.db` and resumes independently. The name is cleaned up for use as a filename, so
spaces and slashes are fine. Reusing a name with different settings still stops, which is the point:
`job_name` separates jobs, it does not switch the check off.

Other options: `disaggregation=True` for category option combos, `ou_attributes=False` to skip the
org-unit group-set columns, and `output_dir` to move the job database.

## How the requests are split up

A download of *n* data items over *m* periods can be sliced into requests in several ways, and
`download_data` picks one automatically from how many of each you asked for. You can override it:

```python
df = dfetch.download_data(..., download_strategy='by_period')
```

| Strategy | One request per … |
|---|---|
| `by_uid` | data item, covering all periods |
| `by_period` | period, covering all data items |
| `batched_uids` | period × batch of 20 data items |
| `batched_periods` | data item × batch of 20 periods |
| `auto` (default) | chosen from the size of the job |

**Worth trying if a download misbehaves.** On some instances a query that comes back empty, times out, or
errors will succeed when sliced a different way — typically a smaller strategy such as `by_uid` or
`batched_periods`. The causes sit on the server: request size and timeout
limits, memory pressure while a large analytics query is assembled, proxies cutting off slow responses,
and the state of the analytics tables all play a part, and they differ between instances. There is no
reliable way to predict which shape an instance will accept, so if results look wrong or a job keeps
failing, changing `download_strategy` is a quick and easy thing to try before assuming the data isn't there.

The log line at the start of a download tells you what was chosen and how many requests it implies:

```
Download strategy: batched_uids (auto, score 40×24=960) — 24 periods × batches of 20 UIDs (~48 requests)
Download strategy: by_uid (auto, score 2×12=24) — 2 requests (all 12 periods per UID)
```

## Saving your results

The package returns DataFrames but does not write a data file for you, so saving is one line of pandas:

```python
df = dfetch.download_data(...)

df.to_csv('outputs/malaria_2024.csv', index=False)
df.to_excel('outputs/malaria_2024.xlsx', index=False)
df.to_parquet('outputs/malaria_2024.parquet')
```

One caveat: the run metadata on `df.attrs` (`data_available`, `db_path`, `job_name`, `created_date`) is
**lost by `to_csv` and `to_excel`**, which store only the table itself. Read what you need from `.attrs`
before saving, or use `df.to_pickle(...)`, which preserves it:

```python
failed = df.attrs['data_available']['Failed']   # check this before you save
df.to_pickle('outputs/malaria_2024.pkl')        # keeps .attrs intact
```

## Errors

All exceptions inherit from `Dhis2Error`, so `except Dhis2Error` catches anything the package raises:
`AuthenticationError`, `Dhis2ConnectionError`, `Dhis2TimeoutError`, `Dhis2RequestError`,
`JobConflictError`, `NotConnectedError`. Bad arguments raise plain `ValueError`.

Individual failed requests inside a long download do not abort the job — they are recorded and reported
in `df.attrs['data_available']['Failed']`, and retried when you re-run.

## Windows sleep prevention

On Windows, downloads keep the machine awake for their duration and release it when finished, so a long
job doesn't stall when the laptop suspends. This is automatic and needs nothing from you. The display is
unaffected, and it is a no-op on other platforms. To wrap your own long-running cell the same way:

```python
from dhis2_fetch import keep_awake

with keep_awake():
    ...
```

## Development

```bash
pip install -e ".[dev]"
pytest
```

The test suite is offline, no DHIS2 server needed. CI runs it on Python 3.9 through 3.13.

`examples/smoke_test.py` is a separate, manual end-to-end check against a live server. It defaults to the
public DHIS2 demo (`admin`/`district`) and is configurable by environment variable:

```bash
python examples/smoke_test.py

DHIS2_URL=https://your-instance.org DHIS2_USERNAME=you DHIS2_PASSWORD=... DHIS2_INDICATOR_UID=... DHIS2_DATASET_UID=... DHIS2_LEVEL=4 python examples/smoke_test.py
```

The public demo servers often have no analytics tables built; when that happens, the indicator and
reporting-metric sections report SKIPPED, and you can point it at an instance with analytics to exercise
those paths.

## Requirements

Python 3.9+ - `requests` - `pandas` - `tqdm`

Tested on Python 3.9 through 3.13, on Linux and Windows. The public functions carry type hints.

## API stability

The package works and is tested, but the version is still `0.x`: the functions and their arguments may
still change between minor versions while the interface settles. Pin a version if you are building
something that has to keep working unattended. Changes are listed in [CHANGELOG.md](https://github.com/Valerian8/dhis2-fetch/blob/main/CHANGELOG.md).

## Licence

MIT — see [LICENSE](https://github.com/Valerian8/dhis2-fetch/blob/main/LICENSE).

## Disclaimer

An independent tool, not affiliated with or endorsed by the DHIS2 project, the University of Oslo, or any
ministry of health. DHIS2 is a trademark of the University of Oslo.
