Metadata-Version: 2.5
Name: lgopy-catalog
Version: 2.0.0
Summary: Reusable block catalog for LgoPy pipelines.
License-Expression: Apache-2.0
License-File: LICENSE
Requires-Python: >=3.10
Requires-Dist: fsspec>=2024.10.0
Requires-Dist: lgopy==2.0.0
Requires-Dist: numpy>=1.26.0
Provides-Extra: rag
Requires-Dist: asyncpg>=0.31.0; extra == 'rag'
Requires-Dist: google-genai>=1.0.0; extra == 'rag'
Requires-Dist: pgvector>=0.4.2; extra == 'rag'
Requires-Dist: sqlalchemy[asyncio]>=2.0.48; extra == 'rag'
Provides-Extra: rag-db
Requires-Dist: asyncpg>=0.31.0; extra == 'rag-db'
Requires-Dist: pgvector>=0.4.2; extra == 'rag-db'
Requires-Dist: sqlalchemy[asyncio]>=2.0.48; extra == 'rag-db'
Provides-Extra: rag-gemini
Requires-Dist: google-genai>=1.0.0; extra == 'rag-gemini'
Description-Content-Type: text/markdown

# lgopy-catalog

Reusable block catalog for LgoPy pipelines.

## Installation

```bash
pip install lgopy-catalog
```

For local workspace development, run from the repository root:

```bash
uv sync
```

## Usage

`lgopy-catalog` works with block packages generated by `Block.build()`:

```text
blocks/
  normalize/
    1.0.0/
      block.py
      __init__.py
      requirements.txt
      manifest.json
      schema.json
```

Publish an existing package directory into a catalog:

```python
from lgopy_catalog import BlockCatalog, FSSpecBlockStore

catalog = BlockCatalog(block_store=FSSpecBlockStore("gs://my-lgopy-catalog"))
catalog.publish_package(".lgopy/blocks/normalize/1.0.0")
```

Search and load packages:

```python
matches = catalog.search("normalization")
Normalize = catalog.load("normalize", "1.0.0")
block = catalog.create("normalize", version="1.0.0", scale=10.0)
```

## Build a pipeline from published blocks

Install `lgopy` in the execution environment and install the selected packages'
`requirements.txt` files; the catalog does not install them automatically.
For a published `normalize` version `1.0.0` whose constructor accepts `scale`:

```python
from lgopy.core import LgoPipeline

steps = [
    {"block": "normalize", "version": "1.0.0", "args": {"scale": 10.0}},
]
report = LgoPipeline.validate(steps, catalog=catalog)
if not report["valid"]:
    raise ValueError(report["issues"])
pipeline = LgoPipeline.from_list(steps, catalog=catalog)
# result = pipeline(your_inputs)
pipeline.save("pipeline.json")
restored = LgoPipeline.from_file("pipeline.json", catalog=catalog)
```

No import of the publisher's block module is needed. See the
[complete publishing and pipeline tutorial](../../mkdocs/guide/catalog-pipelines.md)
for runnable examples and dependency preparation.

## Semantic search

Install the optional semantic-search dependencies:

```bash
pip install "lgopy-catalog[rag]"
```

Configure PostgreSQL and the embedding model with environment variables:

```bash
export LGOPY_CATALOG_DB_HOST=localhost
export LGOPY_CATALOG_DB_PORT=5432
export LGOPY_CATALOG_DB_NAME=lgopy_catalog
export LGOPY_CATALOG_DB_USER=postgres
export LGOPY_CATALOG_DB_PASSWORD=postgres
export LGOPY_CATALOG_EMBEDDING_DIM=768
```

Pass an embedding adapter to enable semantic indexing and search:

```python
from lgopy_catalog import BlockCatalog, FSSpecBlockStore, GeminiEmbedding

catalog = BlockCatalog(
    block_store=FSSpecBlockStore("gs://my-lgopy-catalog"),
    embeddings=GeminiEmbedding(model_id="gemini-embedding-001"),
)

catalog.publish_package(".lgopy/blocks/normalize/1.0.0")
matches = catalog.semantic_search("normalize multispectral imagery before NDVI", k=5)
```

Published block embeddings are generated from manifest metadata, schema details,
and the block class `call` method signature, docstring, and source extracted from
`block.py`.

The catalog package uses a compact layout:

```text
lgopy_catalog.catalog      # BlockCatalog API
lgopy_catalog.store        # block package stores
lgopy_catalog.models       # small dataclasses and protocols
lgopy_catalog.schemas      # database table schemas
lgopy_catalog.utils        # manifest, call-method, runtime helpers
lgopy_catalog.rag          # embeddings and vector search
```

Built-in embedding adapters:

```text
GeminiEmbedding
OllamaEmbedding
```

Use Ollama instead of Gemini:

```python
from lgopy_catalog import BlockCatalog, FSSpecBlockStore, OllamaEmbedding

catalog = BlockCatalog(
    block_store=FSSpecBlockStore("gs://my-lgopy-catalog"),
    embeddings=OllamaEmbedding(model_id="embeddinggemma"),
)

matches = catalog.semantic_search(
    "block for vegetation index calculation",
)
```

Gemini configuration:

```bash
export GOOGLE_API_KEY=...
export LGOPY_CATALOG_GEMINI_EMBEDDING_MODEL_ID=gemini-embedding-001
export LGOPY_CATALOG_GEMINI_OUTPUT_DIM=768
```

Ollama configuration:

```bash
export LGOPY_CATALOG_OLLAMA_BASE_URL=http://localhost:11434
export LGOPY_CATALOG_OLLAMA_EMBEDDING_MODEL_ID=embeddinggemma
```

You can also provide your own embedding adapter by implementing the
`EmbeddingModel` protocol:

```python
import numpy as np

class CustomEmbedding:
    model_name = "custom"

    async def embed(self, text: str) -> np.ndarray:
        return np.array(my_embedding_function(text), dtype=np.float32)

catalog = BlockCatalog(
    block_store=FSSpecBlockStore("gs://my-lgopy-catalog"),
    embeddings=CustomEmbedding(),
)
```

Remove a package version:

```python
catalog.remove_package("normalize", "1.0.0")
```

`catalog.search()` performs text matching with metadata filters; semantic search
ranks indexed candidates by meaning. Each semantic result includes a name,
version, manifest, schema, model name, and cosine distance (lower is closer).
Use `distance_threshold` for an optional maximum distance, not a confidence score.
Inspect a candidate's schema and validate its arguments before execution.

The built-in index needs a reachable PostgreSQL server with pgvector and suitable
initialization permissions. Existing packages are not automatically indexed when
embeddings are enabled: republish them through the configured catalog. Query and
index embeddings must use the same model and vector dimension. Changing models
requires indexing packages for that model; give distinct configurations distinct
`model_name` values. See the [semantic-search guide](../../mkdocs/guide/semantic-search.md)
for complete setup, indexing, and result-to-pipeline examples.
