Metadata-Version: 2.5
Name: bilbysearch
Version: 0.2.0
Summary: Article embeddings with intfloat/multilingual-e5-large-instruct.
Project-URL: Homepage, https://github.com/bilbyai/bilbysearch
Project-URL: Repository, https://github.com/bilbyai/bilbysearch
Author: Samson
License-Expression: MIT
License-File: LICENSE
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.12
Requires-Dist: fsspec
Requires-Dist: gcsfs
Requires-Dist: google-cloud-storage
Requires-Dist: numpy
Requires-Dist: pandas
Requires-Dist: pyarrow
Requires-Dist: sentence-transformers
Requires-Dist: torch
Provides-Extra: all
Provides-Extra: embed
Provides-Extra: gcs
Description-Content-Type: text/markdown

# bilbysearch

**bilbysearch** turns articles into vectors with
[intfloat/multilingual-e5-large-instruct](https://huggingface.co/intfloat/multilingual-e5-large-instruct),
and reads and writes the daily parquet files those vectors live in.

It exists so the model's contract lives in one place. The model id, the 1024-d output
width, the instruction prefix queries need and the `title+body` join rule are all in
[`config.py`](bilbysearch/config.py), so notebooks, batch backfills and retrieval
pipelines cannot drift apart.

## Install

```bash
pip install bilbysearch
```

That one command brings everything: the model stack (sentence-transformers, torch) and
`gs://` support (fsspec, gcsfs, google-cloud-storage). torch is the heavy part, at around
2 GB once CUDA wheels are involved.

The `[embed]`, `[gcs]` and `[all]` extras from 0.1.x are still accepted, so an existing
`"bilbysearch[embed,gcs]"` dependency keeps working, but they no longer add anything.


## Usage

```python
import bilbysearch

model = bilbysearch.load_embedder()

# Embed a frame of articles. Rows with no text get a null vector rather than a
# meaningless one, and three provenance columns record how the vectors were made.
articles = bilbysearch.read_parquet("daily_2026-01-31.parquet")
embedded = bilbysearch.embed_dataframe(articles, model=model)
bilbysearch.write_parquet(embedded, "daily_2026-01-31_embedding.parquet")
```

Queries and passages are embedded by **different functions on purpose**: this model
expects an instruction prefix on queries and none on passages. Skipping it measurably
degrades retrieval.

```python
query = bilbysearch.embed_query("Chinese monetary policy easing", model=model)
passages = bilbysearch.embed_texts(["The PBoC cut the reserve requirement ratio."], model=model)
similarity = query @ passages[0]  # vectors are L2-normalised, so this is cosine
```

Every function other than `load_embedder` works on any object exposing
sentence-transformers' `encode`, so a caller that already has a model can pass it
straight in.

## Development

```bash
uv sync
uv run pytest
uv run ruff check . && uv run ruff format .
```

The tests use a stub model, so the suite runs in well under a second and never downloads
the real 2.2 GB one.

## Licence

MIT — see [LICENSE](LICENSE).
