Metadata-Version: 2.5
Name: bookshelf
Version: 1.0.0rc3
Summary: Python SDK for the Bookshelf data platform
Project-URL: Homepage, https://github.com/climate-resource/bookshelf
Project-URL: Documentation, https://climate-resource.github.io/bookshelf/latest/
Project-URL: Source, https://github.com/climate-resource/bookshelf
Project-URL: Changelog, https://climate-resource.github.io/bookshelf/latest/changelog/
Project-URL: Issues, https://github.com/climate-resource/bookshelf/issues
Author: Climate Resource
License-Expression: MIT
License-File: LICENSE
Classifier: Development Status :: 5 - Production/Stable
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Atmospheric Science
Classifier: Typing :: Typed
Requires-Python: >=3.12
Requires-Dist: filelock>=3.15
Requires-Dist: httpx<1,>=0.27
Requires-Dist: pandas>=2.3
Requires-Dist: platformdirs>=4
Requires-Dist: pyarrow>=18
Requires-Dist: pydantic>=2.7
Requires-Dist: pyyaml>=6.0.2
Requires-Dist: typer>=0.26
Provides-Extra: figures
Requires-Dist: matplotlib>=3.10; extra == 'figures'
Requires-Dist: seaborn>=0.13; extra == 'figures'
Provides-Extra: publish
Requires-Dist: gitpython>=3.1; extra == 'publish'
Requires-Dist: nbconvert>=7.3; extra == 'publish'
Requires-Dist: nbformat>=5.1; extra == 'publish'
Provides-Extra: scmrun
Requires-Dist: scmdata>=0.18; extra == 'scmrun'
Description-Content-Type: text/markdown

# Bookshelf Python SDK

The `bookshelf` package is the official Python SDK for the Bookshelf data platform.
It provides a facade for consuming published data,
producing managed resources,
and running record and replay publishing workflows.
It also includes the `bookshelf` command line interface for authentication,
discovery,
and local cache management.

- [Documentation](https://climate-resource.github.io/bookshelf/latest/)
- [Migrating from 0.4](https://climate-resource.github.io/bookshelf/latest/migrating/)
- [Changelog](https://climate-resource.github.io/bookshelf/latest/changelog/)

## Installation

Install the SDK from PyPI:

```bash
uv add bookshelf
```

The SDK requires Python 3.12 or newer.
It ships with pandas and PyArrow, so `as_df()` and `as_arrow()` work out of the box.
`as_polars()` uses Polars if you have it installed.

Install optional integrations as needed:

```bash
uv add "bookshelf[scmrun,publish]"
```

## Consuming published data

The `bookshelf` package provides the `Bookshelf` facade.
Book coordinates resolve the latest published edition unless `edition=` pins one.
Indexing a `Book` returns a `BookEntry` with book scoped exploration helpers.

`search_volumes()` finds volumes by free text plus discovery filters,
and `volume()` resolves one, carrying the versions and editions it has published.
Both are available on either facade, and neither needs credentials for public data.

```python
from bookshelf import Bookshelf

with Bookshelf() as bs:
    found = bs.search_volumes("emissions", deprecated=False)
    volume = bs.volume("primap-hist")
    versions, latest = volume.versions, volume.latest
```

```python
with Bookshelf() as bs:
    entry = bs.book("rcmip-emissions", "v5.1.0")["magicc"]
    frame = entry.as_df(year_min=2020, year_max=2100, filters={"region": "World"})
    sample = entry.preview(year_min=2020, top_n=5, drop_constant=True)
    facets = entry.facets()
```

`as_df()` returns every selected row as pandas, using wide indexed form for timeseries resources.
The converter family also includes `as_long_df()`, `as_scmrun()`, `as_polars()`, and `as_arrow()`.
They all take the same selection:

- `filters` maps a column to a value, or to a list of alternative values.
- `year_min` and `year_max` bound an inclusive year window on a timeseries.
- `server_side=True` has the platform select, rather than the verified cached file.

`as_df()`, `as_polars()` and `as_arrow()` also take `order`,
a list of columns to sort by, each prefixed with `-` to sort descending.

The cached download is shared, so a second conversion costs no network work.
An external pointer has no cached file, so the platform selects it on every call.
An unknown filter column raises `SelectionError`, which is also a `KeyError`, on either route.

`preview()` returns a bounded `DataPreview` instead.
It adds `limit`, `top_n` and `drop_constant`, and takes `order` on dimension columns.
Its `completeness` says whether the sample holds every selected row.

Use `bs.resource(tracking_id)` for an exact machine or provenance path.
`fetch()` verifies the declared SHA256 before storing bytes in the local content cache.
`download(destination)` copies the verified file to a path you own.
`as_path()` returns the cached file itself, which the cache may evict later.

The SDK is synchronous.
From async code, run each call in a worker thread with `asyncio.to_thread`:

```python
import asyncio

book = await asyncio.to_thread(bs.book, "rcmip-emissions", "v5.1.0", edition=2)
frame = await asyncio.to_thread(book["magicc"].as_df)
```

## Producing and curating data

Managed resources are produced only inside an activity.
The activity derives a stable config hash, records runtime provenance, materialises the object,
and sends explicit Usage and Generation lineage to the API.
Bare strings and UUIDs in `used=` are tracking ids.
Use `Used(name=...)` to resolve an input by the name another resource in the same request was given.
Pass `role="plan"` to `register` for the plan the activity followed, such as a method card.
A plan needs a `name`, takes no `used=`, and is linked to every output of the activity rather than derived from its inputs.

```python
from bookshelf import Bookshelf, Used

with Bookshelf() as bs:
    source = bs.book("rcmip-emissions", "v5.1.0")["magicc"]
    with bs.activity(code_ref="github.com/example/model@abc123", config={"scenario": "ssp245"}) as activity:
        output = activity.register(
            transform(source.as_df()),
            type="timeseries",
            name="model/ssp245/output",
            used=[source, Used(name="model/constants")],
        )

    draft = bs.draft_book("model-results", version="v1.0.0")
    draft.attach(output, name_in_book="ssp245")
    draft.publish()
```

`register_external()` is available on both `Bookshelf` and an activity.
The former catalogues an existing pointer.
The latter attributes an external output to the current run.
Book drafting, attachment, and publication remain separate editorial calls.

Use `activity.register_many()` with `RegisterItem` values for a batch.
An atomic batch over 1000 items raises before any upload begins.
A larger non atomic batch is split into requests of at most 1000 items.
If any item fails, the facade finishes every chunk and raises `PartialRegistrationError`.
The error retains indexed successful outcomes, usable committed resource handles,
and each failed index with its typed `ItemError`.
Index `-1` identifies a batch level lineage failure reported by the server.
The server decides:
a registration aliases onto a resource your organisation already holds with the same bytes,
unless a bundle pins its tracking id.
The first canonical resource keeps its name, even when later items supply a different one.
Returned producer handles expose `registration_status` and `registration_outcome`,
so callers can detect this `aliased` result.

A failed multipart PUT can leave an unfinished upload because the server has no abort endpoint.
Registration does not begin after that failure.
A retry reuses the content addressed upload path and safely resumes the workflow.

## Publishing a recorded bundle

`bookshelf record` runs a build file offline and writes a bundle,
and `bookshelf publish` replays that bundle to the platform.
`bookshelf.publisher.replay_bundle` does the same from Python.

A replay uploads the managed bytes,
then sends the whole bundle to `POST /v1/bundles/replay` as one transactional request.
The server registers the resources, mints the recorded activity's provenance edges,
drafts the book, attaches every entry and publishes it,
and rolls all of it back on a failure anywhere.

Every resource travels under its bundle-local name.
A resource that a book entry names takes that name inside the book,
and `used=` lineage cites the name of a resource recorded earlier in the same bundle.
An input the platform already holds is cited by its digest instead.
The server computes the seal from the request,
so replaying the same bundle again converges on the one edition
and reports `converged` rather than minting a rival.

## Uploading a file that cannot be checked in

`bookshelf upload FILE --type TYPE` puts a file on the bookshelf as an input that belongs to no book,
and prints the `bookshelf://sha256/<hex>` URI a recipe declares it by.
`Bookshelf.register_file` does the same from Python,
and `Bookshelf.resource_by_hash` resolves the digest back into the resource.

The command leaves the file hidden, so it is readable by the uploading organisation alone.

## Credentials

`Bookshelf` authenticates through the credential providers in `bookshelf.auth`.
Each is an `httpx.Auth`, so one provider object can be shared between clients:

- `StaticToken`: a fixed bearer token, no refresh.
- `RefreshTokenExchange`: a WorkOS user access/refresh pair from `bookshelf auth login`.
  The refresh token rotates on each use and an `on_rotate` callback persists the new pair.
- `ClientCredentials`: an OAuth2 `client_credentials` machine credential.
  A refresh is a plain re-POST, there is nothing to persist.
- `ActionsOidcToken`: the running job's GitHub Actions OIDC token, minted for the read audience.
  It mints on first use and mints again once the API refuses the one it holds.

Refresh mechanics are shared: proactive refresh five minutes before expiry,
one refresh-and-replay after an unexpected 401 (a second 401 raises `AuthenticationError`),
and single-flight refresh behind one lock, so threads sharing a provider refresh once.
There is no background refresh task.
A token handed in with no known expiry is refreshed before first use,
because it may already be dead server-side.

When `auth=` is omitted, ambient credentials resolve in this order
(explicit beats ambient, machine beats human):

1. `$BOOKSHELF_TOKEN` as a static bearer
2. a GitHub Actions OIDC token, only when `$BOOKSHELF_AUTH=github-actions`
3. `$BOOKSHELF_CLIENT_ID` + `$BOOKSHELF_CLIENT_SECRET`, minted at `$BOOKSHELF_TOKEN_URL`
4. stored `bookshelf auth login` credentials
5. unauthenticated (public reads)

`auth=` also accepts a provider instance or a bare token string,
and an explicit `auth=None` stays unauthenticated.
`base_url` resolves as argument, then `$BOOKSHELF_URL`, then the production URL.

### Client lifecycle in an embedded service

The client is long-lived by design: token state lives in the provider
and each client pools connections.
Construct one client at startup, inject it as a dependency, and close it at shutdown.
In FastAPI that is a lifespan.
A plain `def` endpoint runs in FastAPI's thread pool, so it can call the client directly:

```python
from contextlib import asynccontextmanager

from fastapi import FastAPI, Request

from bookshelf import Bookshelf
from bookshelf.auth import ClientCredentials

@asynccontextmanager
async def lifespan(app: FastAPI):
    with Bookshelf(
        auth=ClientCredentials(client_id, client_secret, token_url=token_url),
    ) as bs:
        app.state.bookshelf = bs
        yield

app = FastAPI(lifespan=lifespan)

@app.get("/co2")
def co2(request: Request):
    bs: Bookshelf = request.app.state.bookshelf
    book = bs.book("rcmip-emissions", "v5.1.0", edition=2)
    frame = book["magicc"].as_long_df(filters={"variable": "Emissions|CO2"})
    return frame.to_dict(orient="records")
```

Do **not** open a client per request (`with Bookshelf(...)` inside a handler):
that churns the connection pool and discards the cached token on every call.
Context managers are optional.
Notebooks can construct a client plainly and never close it.
