Metadata-Version: 2.5
Name: relay-metis
Version: 0.2.0
Summary: Lightweight Python client for Metis, Relay's Cube-backed metrics layer.
License-Expression: MIT
License-File: LICENSE
Requires-Python: >=3.11
Requires-Dist: pandas>=2.0
Requires-Dist: requests>=2.31
Description-Content-Type: text/markdown

# metis

Lightweight Python client for Metis, Relay's Cube-backed metrics layer. Queries a Metis
view and returns a typed, annotated pandas DataFrame — the metrics you see in dashboards,
in a notebook, computed the same way.

## Setup

```bash
export METIS_API_URL="https://<deployment>.cubecloudapp.dev"   # /cubejs-api/v1 optional
export METIS_API_KEY="<pre-minted Cube API token>"
```

Both can also be passed explicitly: `MetisClient(url=..., api_key=...)` (constructor
arguments win over the environment).

## Usage

```python
from datetime import date
from metis import MetisClient

client = MetisClient()

df = client.query(
    "orders",
    measures=["order_count", "revenue_sum"],
    dimensions=["status"],
    time_dimension="created_at",
    granularity="day",
    date_range=(date(2026, 8, 1), date(2026, 8, 31)),
    segments=["is_completed"],
    filters={"is_priority": True},
    order={"created_at": "asc"},
    timezone="Europe/London",
)
```

Member names (measures, dimensions, segments) are unprefixed and scoped to the view;
result columns come back prefix-stripped. Discovery:

```python
client.meta()  # one row per view/member: kind, type, title, description
```

Unknown views/members fail before the query is sent, with did-you-mean suggestions.

### Time dimensions

One time dimension takes the scalar kwargs; `granularity` and `date_range` are both
optional (a range without a granularity filters, a granularity without a range groups
over all time):

```python
time_dimension="created_at", granularity="day", date_range=(date(2026, 8, 1), date(2026, 8, 31))
```

Several take `time_dimensions`, a list of entries with the keys `dimension`,
`granularity` and `date_range`. The scalar form is sugar for a one-entry list, and the
two forms cannot be combined:

```python
time_dimensions = [
    {
        "dimension": "delivered_on",
        "granularity": "day",
        "date_range": ("2026-08-01", "2026-08-31"),
    },
    {
        "dimension": "created_at",
        "date_range": ("2026-07-25", "2026-08-31"),
    },  # filter only
    {"dimension": "due_at", "granularity": "hour"},
]
```

Each dimension may appear once. Every granular entry comes back as its own
`datetime64` column under the bare member name.

`timezone` is a [tz database](https://en.wikipedia.org/wiki/Tz_database) name such as
`"Europe/London"`, validated locally, and applies to bucketing and date ranges. Time
columns stay tz-naive and mean wall-clock time in that zone.

### Filters

The mapping form is sugar for equals/in (`None` means "is not set"):

```python
filters = {"status": ["shipped", "delivered"], "is_priority": True}
```

Anything richer takes raw Cube filter dicts (member names still unprefixed):

```python
filters = [{"member": "revenue_sum", "operator": "gt", "values": ["1000"]}]
```

### Dtypes

Model-driven, never inferred from whichever values a result happens to contain:

| Member | dtype |
| --- | --- |
| count / count_distinct measures | `Int64` (nullable) |
| all other numeric measures | `float64` |
| numeric dimensions, integer-typed at source | `Int64` (nullable) |
| numeric dimensions, decimal-typed at source | `float64` |
| booleans | `boolean` (nullable) |
| time members | `datetime64` |

Numeric dimensions are classified from the raw serialisation: integer columns always
arrive as bare digit strings and classify correctly. Known caveat: a decimal-typed
dimension (FLOAT64 or NUMERIC) whose result values are all whole also arrives as bare
digits and reads as `Int64` until a decimal value appears. Override any column
per-call: `dtypes={"weight_grams": "float64"}`.

Query provenance (payload, Cube annotation, request time) is attached under
`df.attrs["metis"]`.

### Cache mode

`cache` defaults to `"must-revalidate"`: serves cached data while current, waits for
fresh data when expired, never returns stale data. The other Cube modes
(`"no-cache"`, `"stale-while-revalidate"`, `"stale-if-slow"`) are accepted and
validated locally. Prefer `must-revalidate`: `no-cache` skips the refresh-key check,
so each Continue-wait poll schedules a fresh database query instead of attaching to
the in-flight one.

### Escape hatches

```python
client.load(
    {"query": {...}, "cache": "must-revalidate"}
)  # raw Cube REST payload -> raw dict
client.raw_meta()  # raw /meta cubes list, every block Cube emits, including member["meta"]
```

`meta()` projects each member to scalar columns; `raw_meta()` is for anything that
lives in a member's `meta` block.

## Behaviour notes

- Long-running queries are polled (Cube "Continue wait") within a wall-clock budget of
  120s by default — raise `total_wait_budget_seconds` for heavy queries.
- Transient errors retry up to 3 times with backoff; auth failures and query errors
  fail immediately and loudly.
