Metadata-Version: 2.5
Name: bilbysearch
Version: 0.1.0
Summary: Article embeddings with intfloat/multilingual-e5-large-instruct.
Project-URL: Homepage, https://github.com/bilbyai/bilbysearch
Project-URL: Repository, https://github.com/bilbyai/bilbysearch
Author: Samson
License-Expression: MIT
License-File: LICENSE
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.12
Requires-Dist: numpy
Requires-Dist: pandas
Requires-Dist: pyarrow
Provides-Extra: all
Requires-Dist: fsspec; extra == 'all'
Requires-Dist: gcsfs; extra == 'all'
Requires-Dist: google-cloud-storage; extra == 'all'
Requires-Dist: sentence-transformers; extra == 'all'
Requires-Dist: torch; extra == 'all'
Provides-Extra: embed
Requires-Dist: sentence-transformers; extra == 'embed'
Requires-Dist: torch; extra == 'embed'
Provides-Extra: gcs
Requires-Dist: fsspec; extra == 'gcs'
Requires-Dist: gcsfs; extra == 'gcs'
Requires-Dist: google-cloud-storage; extra == 'gcs'
Description-Content-Type: text/markdown

# bilbysearch

**bilbysearch** turns articles into vectors with
[intfloat/multilingual-e5-large-instruct](https://huggingface.co/intfloat/multilingual-e5-large-instruct),
and reads and writes the daily parquet files those vectors live in.

It exists so the model's contract lives in one place. The model id, the 1024-d output
width, the instruction prefix queries need and the `title+body` join rule are all in
[`config.py`](bilbysearch/config.py), so notebooks, batch backfills and retrieval
pipelines cannot drift apart.

## Install

```bash
pip install bilbysearch[embed]
```

The base install is deliberately light — just pandas and pyarrow — so a caller that only
reads the parquet files, or only reads the model contract, does not pull a multi-gigabyte
ML stack it will never use.

| Extra | Adds | Needed for |
|---|---|---|
| *(none)* | pandas, pyarrow, numpy | reading `config`, local parquet IO |
| `[embed]` | sentence-transformers, torch | `load_embedder` |
| `[gcs]` | fsspec, gcsfs, google-cloud-storage | `gs://` URIs |
| `[all]` | both of the above | everything |

Until the first PyPI release, install from the repo instead:

```bash
uv add "bilbysearch[all] @ git+https://github.com/bilbyai/bilbysearch.git"
```

## Usage

```python
import bilbysearch

model = bilbysearch.load_embedder()

# Embed a frame of articles. Rows with no text get a null vector rather than a
# meaningless one, and three provenance columns record how the vectors were made.
articles = bilbysearch.read_parquet("daily_2026-01-31.parquet")
embedded = bilbysearch.embed_dataframe(articles, model=model)
bilbysearch.write_parquet(embedded, "daily_2026-01-31_embedding.parquet")
```

Queries and passages are embedded by **different functions on purpose**: this model
expects an instruction prefix on queries and none on passages. Skipping it measurably
degrades retrieval.

```python
query = bilbysearch.embed_query("Chinese monetary policy easing", model=model)
passages = bilbysearch.embed_texts(["The PBoC cut the reserve requirement ratio."], model=model)
similarity = query @ passages[0]  # vectors are L2-normalised, so this is cosine
```

Only `load_embedder` needs the `[embed]` extra. Every other function works on any object
exposing sentence-transformers' `encode`, so a caller that already has a model can pass
it straight in.

## Development

```bash
uv sync
uv run pytest
uv run ruff check . && uv run ruff format .
```

The tests use a stub model, so the suite runs in well under a second and never downloads
the real 2.2 GB one.

## Licence

MIT — see [LICENSE](LICENSE).
