Metadata-Version: 2.3
Name: gwenlake
Version: 0.8.2
Summary: Gwenlake Python Client
Requires-Dist: click>=8.1.7
Requires-Dist: pyarrow>=14.0
Requires-Dist: httpx>=0.28.1
Requires-Dist: pandas>=3.0.3
Requires-Dist: pydantic>=2.12.5
Requires-Dist: python-dotenv>=1.2.1
Requires-Python: >=3.12
Description-Content-Type: text/markdown

# Gwenlake Python Library

The Gwenlake Python library provides convenient access to the Gwenlake API
from applications written in Python. A single `Gwenlake` client gives you access
to your catalog — projects, datasets, files and SQL.

## Installation

```sh
pip install -U gwenlake
```

Or install the latest development version straight from GitHub:

```sh
pip install -U git+https://github.com/gwenlake/gwenlake-python
```

## Authentication

The client authenticates with a Bearer token, resolved in this order:

1. an explicit `api_key` / `credentials` passed to the client,
2. a named `profile`,
3. the `GWENLAKE_API_KEY` environment variable,
4. the `default` profile in `~/.gwenlake/credentials`.

```bash
export GWENLAKE_API_KEY='sk-...'
```

```python
from gwenlake import Gwenlake

# uses GWENLAKE_API_KEY, or the default ~/.gwenlake/credentials profile
client = Gwenlake()

# or pass the key explicitly
client = Gwenlake(api_key="sk-...")

# or pick a profile from ~/.gwenlake/credentials
client = Gwenlake(profile="myteam")
```

The `~/.gwenlake/credentials` file is an INI file with one section per profile,
holding either a static `token` (API key) or OAuth2 `client_id` / `client_secret`.

## Projects

```python
projects = client.projects.list()
for p in projects:
    print(p["alias"], p["id"])

project = client.projects.get("res.project.…")
```

## Datasets

```python
datasets = client.datasets.list()
for d in datasets:
    print(d["alias"], d["id"])

dataset = client.datasets.get("res.dataset.…")
```

## Files

Files live inside a dataset.

```python
dataset_id = "res.dataset.…"

# list files
for f in client.files.list(dataset_id):
    print(f["filename"], f["file_size"])

# upload a local file (optionally into a subdirectory with path=...)
client.files.upload(dataset_id, "report.pdf")
client.files.upload(dataset_id, "report.pdf", path="docs")

# download a file
content = client.files.download(dataset_id, "report.pdf")

# presigned URL / delete
url = client.files.presigned_url(dataset_id, "report.pdf")
client.files.delete(dataset_id, "report.pdf")
```

## SQL

Run SQL against a dataset (DuckDB), referencing it as
`'<project_alias>.<dataset_alias>'`. With `format="json"` the rows are returned
under `data`:

```python
result = client.statements.create(
    statement="SELECT * FROM 'flights.flight-data' LIMIT 10",
    format="json",
)
for row in result["data"]:
    print(row)
```

Pass a `connection_id` to run the statement against a connection's native engine
(PostgreSQL, S3, …) instead of a dataset.

## Transforms

A Palantir Foundry-style transforms layer (`gwenlake.transforms`) lets you write
dataset-to-dataset transformations as decorated functions. Datasets are
addressed as `"<project_alias>.<dataset_alias>"` — the same handle used in SQL.

`transform_df` — the function receives each `Input` as a `pandas.DataFrame` and
**returns** the DataFrame to write to the (single) `Output`. The result is
written automatically (snapshot/replace by default):

```python
from gwenlake.transforms import transform_df, Input, Output

@transform_df(
    raw_data=Input("Project_A.users"),
    processed_data=Output("Project_A.users_filtered"),
)
def process(raw_data):
    df = raw_data[raw_data["age"] >= 18].copy()
    df["name_upper"] = df["name"].str.upper()
    return df

process(client)   # reads, computes, writes
```

`transform` — the lower-level form: the function receives `TransformInput` /
`TransformOutput` objects and reads/writes explicitly. Use it for non-tabular
data (images, PDFs, …) via `.filesystem()`:

```python
from gwenlake.transforms import transform, Input, Output

@transform(
    my_input=Input("Project_A.users"),
    my_output=Output("Project_A.users_distinct"),
)
def dedupe_users(my_input, my_output):
    df = my_input.dataframe()
    # mode="replace" (default) clears the dataset first; "append" keeps existing files
    my_output.write_dataframe(df.drop_duplicates(), mode="replace")

@transform(
    images=Input("Project_A.scans"),
    thumbnails=Output("Project_A.scans_processed"),
)
def process_files(images, thumbnails):
    src, dst = images.filesystem(), thumbnails.filesystem()
    for entry in src.ls():
        data = src.read(entry["filename"])          # raw bytes (PDF, image, …)
        with dst.open(f"copy/{entry['filename']}", "wb") as f:
            f.write(data)
```

**Large datasets** — page through with `LIMIT/OFFSET` instead of loading
everything at once. `iter_dataframes()` yields `pandas.DataFrame` chunks and
`write_dataframes()` streams them back out as `part-00000.parquet`, …:

```python
@transform(
    big_dataset=Input("Project_A.events"),
    result=Output("Project_A.events_clean"),
)
def transform_in_chunks(big_dataset, result):
    chunks = (
        chunk[chunk["valid"]]
        for chunk in big_dataset.iter_dataframes(chunk_size=50_000, order_by="id")
    )
    result.write_dataframes(chunks, mode="replace")
```

Pass `order_by=` for a deterministic page split. The transforms layer is
synchronous.

## Async

Every resource is also available on `AsyncGwenlake`:

```python
import asyncio
from gwenlake import AsyncGwenlake

async def main():
    client = AsyncGwenlake()
    print(await client.projects.list())

asyncio.run(main())
```

See [`examples/`](examples/) for runnable scripts.
