Metadata-Version: 2.4
Name: akshamdb
Version: 0.1.1
Summary: Fast, persistent, embeddable vector database
License: Apache-2.0
Keywords: vector database,semantic search,embeddings,RAG,HNSW,ANN
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: C++
Classifier: Topic :: Database
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: numpy>=1.24
Provides-Extra: server
Requires-Dist: flask>=2.3; extra == "server"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-benchmark; extra == "dev"

# akshamdb

**A fast, persistent, embeddable vector database.**

`akshamdb` runs in your Python process — no server, no Docker, no external
service. You give it vectors, it stores them on disk and gives you similarity
search back. It is built for semantic search and RAG applications where you
want to own the whole stack.

- **Embeddable** — `import akshamdb`, open a directory, done.
- **Persistent** — a crash-safe write-ahead log plus an on-disk index. Reopen
  and your data is still there.
- **Hybrid search** — vector cosine similarity blended with BM25 keyword
  scoring, plus metadata filtering, MMR diversity, query expansion, score
  calibration and pluggable rerankers.
- **Small footprint** — the only required dependency is `numpy`. The nearest-
  neighbour index is a C++ HNSW implementation exposed through pybind11.

---

## Installation

```bash
pip install akshamdb
```

The C++ index extension is compiled during installation, so you need a build
toolchain available:

- Python 3.10+
- A C++ compiler (GCC / Clang / MSVC)
- CMake 3.15+

Optional extras:

```bash
pip install "akshamdb[server]"   # Flask REST API server + akshamdb-server CLI
pip install "akshamdb[dev]"      # pytest + pytest-benchmark
```

`akshamdb` does **not** ship an embedding model. Bring your own vectors from
whichever model you use (Sentence-Transformers, OpenAI, Cohere, …).

---

## Quick start

```python
import numpy as np
import akshamdb

with akshamdb.open("./my_db", dim=768) as db:
    db.add("doc_1", np.random.rand(768), text="hello world",
           metadata={"source": "wiki"})

    db.add_many(
        ids=["doc_2", "doc_3"],
        vectors=[np.random.rand(768), np.random.rand(768)],
        texts=["second doc", "third doc"],
        metadata=[{"source": "blog"}, {"source": "wiki"}],
    )

    results = db.search(np.random.rand(768), top_k=5)
    for r in results:
        print(r["score"], r["id"], r["text"])
# close() is called automatically by the context manager
```

Without the context manager, always call `db.close()` before your process
exits so pending writes and the index are flushed to disk.

---

# Operations reference

Everything below is reachable from the top-level `akshamdb` package.

## Module — `akshamdb`

| Call | Description |
|------|-------------|
| `akshamdb.open(path, *, dim=768)` | Open (or create) a database directory. Returns an `AkshamDB` handle. `dim` is any positive integer — `768` is only the default (common: 384, 768, 1536, 3072). Every vector added to this database must have that dimension. |
| `akshamdb.register_reranker(name, factory)` | Register a reranker under a string name so `reranker="name"` works. `factory` may be a class (new instance per search), a lambda, or a singleton `BaseReranker` instance. `"default"` is reserved. |
| `akshamdb.__version__` | Package version string. |

**Exceptions**

| Class | Meaning |
|-------|---------|
| `akshamdb.AkshamError` | Base class for every error `akshamdb` raises. Catch this to handle any DB error. |
| `akshamdb.DimensionError` | Raised when a vector's length does not match the database `dim`. Subclass of `AkshamError`. |

**Re-exported types** (also importable straight from `akshamdb`):
`SearchPipeline`, `PipelineReranker`, `SearchResult`, `SearchContext`,
`BaseReranker`, `HybridSignalReranker`, `FeatureReranker` (alias of
`HybridSignalReranker`), `CrossEncoderReranker`, `QueryTelemetry`.

---

## Database handle — `AkshamDB`

Returned by `akshamdb.open()`. Do not construct directly. All methods are
thread-safe.

### Writing

| Call | Description |
|------|-------------|
| `db.add(id, vector, *, text="", metadata=None)` | Insert or overwrite one document. `text` feeds the BM25 keyword index; `metadata` is a dict of filterable key/value pairs. |
| `db.add_many(ids, vectors, *, texts=None, metadata=None)` | Batch insert/overwrite. `ids` and `vectors` must be the same length. Much faster than a loop. |
| `db.delete(id)` | Remove a document. No-op if the id is unknown. The vector slot stays in the HNSW graph but is invisible to all searches until the next full `rebuild_index()`. |

### Reading

| Call | Description |
|------|-------------|
| `db.get(id)` | Return the full document dict (including the raw vector under `"vector"`), or `None` if not found. |
| `db.search(vector, ...)` | Similarity search for one query. Returns `list[SearchResult]`, highest score first. Options below. |
| `db.search_batch(vectors, ...)` | Same as `search` for many queries at once, using parallel ANN traversal and batched scoring. Returns `list[list[SearchResult]]`, one list per input vector, in input order. |

### `db.search()` / `db.search_batch()` options

| Argument | Default | Meaning |
|----------|---------|---------|
| `top_k` | `5` | Number of results to return. |
| `filter` (`search`) / `filter` (`search_batch`, applied to every query) | `None` | Metadata filter dict — see **Filter DSL** below. |
| `text` (`search`) / `texts` (`search_batch`) | `""` / `None` | Query text for hybrid BM25 + vector scoring. Quoted substrings act as phrase constraints. |
| `alpha` | `0.7` | Vector weight in `[0, 1]`. `1.0` = pure vector, `0.0` = pure BM25. |
| `mmr` | `False` | Enable Maximal Marginal Relevance result diversification. |
| `mmr_lambda` | `0.5` | Relevance/diversity tradeoff in `[0, 1]`. `1.0` = pure relevance, `0.0` = pure diversity. |
| `expand_query` | `False` | Pseudo-Relevance Feedback: append high-IDF terms from the top BM25 hits to the query before final scoring. Uses only the local BM25 index. |
| `calibrate` | `False` | Normalise final scores. |
| `calibrate_method` | `"softmax"` | `"softmax"` (scores sum to 1), `"minmax"` (scores in `[0, 1]`), or `"zscore"` (sigmoid of z-score). |
| `reranker` | `None` | `None`, `"default"`, a registered name, or a `BaseReranker` instance — see **Rerankers**. |
| `reranker_factor` | `4` | Candidate multiplier when reranking: retrieval fetches `max(top_k * factor, 20)` candidates and the reranker picks the final `top_k`. |
| `pipeline` | `None` | A `SearchPipeline` — its settings win; the explicit arguments above are used only for values the pipeline did not set. |

### Transactions

| Call | Description |
|------|-------------|
| `with db.transaction(): ...` | Atomic batch. All inserts inside commit together on clean exit, or roll back entirely on any exception. Also the fastest way to bulk-load (single WAL flush + single index save). |

### Maintenance & introspection

| Call | Description |
|------|-------------|
| `db.info()` | Dict: `documents`, `index_nodes`, `data_path`, `performance` (per-op timing if metrics enabled). |
| `db.rebuild_index(progress_callback=None)` | Force a full HNSW rebuild from the LSM store. `progress_callback(indexed, total)` is called after each batch. Normally automatic on `open()` when the saved node count is stale. |
| `db.enable_telemetry(log_path=None, log_results=False)` | Turn on structured per-query logging. Returns a `QueryTelemetry` — see below. |
| `db.close()` | Flush pending writes and save the index. Always call this (or use the context manager). |

### Python protocol

| Expression | Description |
|------------|-------------|
| `with akshamdb.open(...) as db:` | Context manager — `db.close()` runs on exit. |
| `len(db)` | Number of documents currently stored. |
| `"doc_1" in db` | Membership test by id. |
| `repr(db)` | `AkshamDB(documents=…, dim=…, path=…)`. |

---

## Filter DSL

Used by `db.search(filter=...)`, `db.search_batch(filter=...)`, and
`SearchPipeline.metadata_filter(...)`. Multiple keys are ANDed.

| Form | Match |
|------|-------|
| `{"key": "value"}` | Exact match (`str` / `int` / `float` / `bool`). |
| `{"key": ["a", "b"]}` | OR — any value in the list. |
| `{"key": {"not": "v"}}` | Negation — every doc whose `key` ≠ `"v"`. |
| `{"key": {"gte": 0, "lte": 9}}` | Numeric range. Operators: `gte`, `gt`, `lte`, `lt` (combine freely). |

A key that was never seen in any document's metadata matches zero results.

---

## Pipeline (fluent builder)

Configure search once, reuse it across queries. Every method returns `self`,
so calls chain. The object is safe to share across `db.search()` calls.

| Call | Description |
|------|-------------|
| `akshamdb.SearchPipeline()` | Create a reusable pipeline (defaults: `alpha=0.7`, everything else off). |
| `.hybrid(alpha=0.7)` | Set the vector/BM25 blend weight. Raises `ValueError` outside `[0, 1]`. |
| `.mmr(lambda_=0.5)` | Enable MMR diversity. `1.0` = pure relevance, `0.0` = pure diversity. |
| `.query_expansion()` | Enable Pseudo-Relevance Feedback query expansion. |
| `.metadata_filter({...})` **or** `.metadata_filter(key="value", ...)` | Attach a metadata filter (full Filter DSL above). |
| `.reranker(r, factor=4)` | Attach a reranker: `"default"`, a registered name, or a `BaseReranker`. `factor` is the candidate multiplier. |
| `.calibrate(method="softmax")` | Normalise scores — `"softmax"` / `"minmax"` / `"zscore"`. Raises `ValueError` on any other value. |

```python
pipe = (akshamdb.SearchPipeline()
        .hybrid(alpha=0.8)
        .query_expansion()
        .metadata_filter(country="Japan")
        .reranker("default")
        .mmr(lambda_=0.6)
        .calibrate())

results = db.search(q_vec, text=query, top_k=10, pipeline=pipe)
batched = db.search_batch(q_vecs, texts=queries, top_k=10, pipeline=pipe)
```

`repr(pipe)` prints the configured stages.

---

## Rerankers

Reranking runs after retrieval on `max(top_k * factor, 20)` candidates.

### How `reranker=` is resolved

| Value | Behaviour |
|-------|-----------|
| `None` | No reranking — retrieval order is returned. |
| `"default"` | Built-in `HybridSignalReranker`, wired to this DB's live indexes. Lazily created and cached per handle. |
| a registered name | Looked up in the registry (`akshamdb.register_reranker(...)`). |
| a `BaseReranker` instance | Used as-is. |

### `BaseReranker` (abstract)

Subclass and implement:

```python
class MyReranker(akshamdb.BaseReranker):
    def rerank(self, query, results, top_k=None):
        # results: list of dicts with "id", "score", "text", "metadata"
        # return them sorted best-first; set "rerank_score" for inspection
        ...
```

Override `rerank_with_context(context, results)` instead when you need the
`SearchContext` (filters, alpha, raw query vector, pipeline). The default
implementation delegates to `rerank()`.

### `HybridSignalReranker(hybrid_search, weights=None, min_proximity_terms=2)`

Built-in default. **Zero external dependencies.** Blends four signals computed
from `akshamdb`'s own indexes (default weights, auto-normalised to sum to 1):

| Signal | Weight | Source |
|--------|--------|--------|
| `vector` | `0.45` | Cosine score from the ANN stage, min-max normalised across candidates. |
| `bm25` | `0.30` | Okapi BM25 (`k1=1.5`, `b=0.75`) from the inverted index. |
| `exact_match` | `0.15` | Fraction of stemmed query terms present in the doc (own Porter stemmer). |
| `proximity` | `0.10` | Tightness of the smallest window covering all query terms (positional index; skipped below `min_proximity_terms` distinct terms). |

Pass `weights={"bm25": 0.4}` to override individual weights (partial dict is
fine). Each result gains `rerank_score`, `_signals` and `_signal_weights`
(used by `SearchResult.explanation()`).

Normally you don't construct it — use `reranker="default"`.

### `FeatureReranker`

Backward-compatible alias of `HybridSignalReranker`.

### `CrossEncoderReranker(model_name="cross-encoder/ms-marco-MiniLM-L-6-v2", device=None, batch_size=8)`

**External integration** — requires `pip install sentence-transformers`.
Wraps `sentence_transformers.CrossEncoder`. Needs a text query. `device` is
`"cpu"` / `"cuda"` / `"mps"` / `None` (auto). The model loads lazily on first
use. `rerank(..., calibrate=True)` (default) passes raw logits through a
sigmoid.

### `PipelineReranker([stage1, stage2, ...])`

Chains rerankers left to right. Intermediate stages see the full candidate set
(no pruning); only the last stage applies `top_k`. Use it to run a cheap pass
before an expensive one:

```python
chained = akshamdb.PipelineReranker([
    akshamdb.HybridSignalReranker(db._engine.query_engine.hybrid_search),
    akshamdb.CrossEncoderReranker(),
])
results = db.search(q_vec, text=query, top_k=5, reranker=chained)
```

### Plugin registry

```python
akshamdb.register_reranker("cohere", CohereReranker)          # class
akshamdb.register_reranker("cohere", lambda: CohereReranker(key=...))  # factory
akshamdb.register_reranker("cohere", CohereReranker(key=...))  # singleton
results = db.search(vec, text=query, reranker="cohere")
```

---

## Result objects — `SearchResult`

A `dict` subclass, so `r["score"]` and every dict operation still work.

| Key (always present) | |
|----------------------|---|
| `id` | Document identifier. |
| `score` | Raw retrieval score (cosine or hybrid). |
| `text` | Document text. |
| `metadata` | Your metadata dict. |

| Key (conditional) | Present when |
|-------------------|--------------|
| `raw_score` | `calibrate=True` — the pre-calibration score. |
| `rerank_score` | A reranker ran. |
| `_signals` / `_signal_weights` | `HybridSignalReranker` ran. |

| Property / method | Description |
|-------------------|-------------|
| `r.id` | `str` id. |
| `r.score` | Effective ranking score: `rerank_score` if reranking ran, else the retrieval score. |
| `r.retrieval_score` | Raw ANN/hybrid score, before any reranking. |
| `r.text`, `r.metadata` | Typed accessors. |
| `r.raw_score` | Pre-calibration score or `None`. |
| `r.rerank_score` | Reranker score or `None`. |
| `r.explanation()` | Human-readable, multi-line breakdown of every scoring stage and each signal's weighted contribution. |

### `SearchContext` (passed to `rerank_with_context`)

Frozen dataclass with: `query`, `vector`, `top_k`, `filters`, `alpha`,
`pipeline`.

---

## Telemetry — `QueryTelemetry`

```python
tel = db.enable_telemetry("./aksham_queries.jsonl", log_results=False)
db.search(vec, text=query)        # logged automatically
print(tel.summary())
```

| Call | Description |
|------|-------------|
| `db.enable_telemetry(log_path=None, log_results=False)` | Enable logging; returns the `QueryTelemetry`. `log_path=None` keeps in-memory stats only. `log_results=True` includes the result list in each record. |
| `tel.summary()` | Rolling stats dict: `total_queries`, `queries_in_buffer`, `avg_latency_ms`, `p50_latency_ms`, `p95_latency_ms`, `avg_results`, `avg_top_score`. |
| `tel.close()` | Close the log file. |

Every `db.search()` / `db.search_batch()` call writes one JSON object per line
with: `timestamp`, `query_text`, `top_k`, `filters`, `alpha`, `mmr`,
`expand_query`, `calibrate`, `num_results`, `top_score`, `latency_ms`,
`stage_ms`, and `results` (when `log_results=True`).

---

## Configuration — `akshamdb.utils.schema.Config`

Engine-wide defaults. Set attributes before `akshamdb.open()`.

| Attribute | Default | Purpose |
|-----------|---------|---------|
| `MEMTABLE_LIMIT` | `1000` | Docs held in memory before flushing to an SSTable. |
| `COMPACT_THRESHOLD` | `4` | Merge SSTables once this many accumulate. |
| `HNSW_M` | `16` | Graph connections per node (higher = better recall, more memory). |
| `HNSW_EF_CONSTRUCTION` | `200` | Build quality (higher = better recall, slower inserts). |
| `HNSW_EF_SEARCH` | `150` | Candidates explored per query (higher = better recall, slower search). |
| `ENABLE_PARALLEL_SEARCH` | `True` | GIL-free parallel scoring. |
| `PARALLEL_SEARCH_WORKERS` | `None` | Thread count; `None` = `os.cpu_count()`. |
| `ENABLE_PQ` | `False` | Product Quantization for memory compression on very large corpora. |
| `PQ_M`, `PQ_K` | `96`, `256` | PQ subspaces / centroids per subspace. |
| `ENABLE_IVF` | `False` | IVF coarse index for 10M+ vector corpora. |
| `IVF_N_CLUSTERS`, `IVF_NPROBE` | `256`, `8` | IVF cell count / cells probed per query. |

---

## REST API (optional)

```bash
pip install "akshamdb[server]"
akshamdb-server                 # serves on :8000, data in ./data
AKSHAMDB_PATH=/my/db AKSHAMDB_DIM=768 AKSHAMDB_PORT=9000 akshamdb-server
```

| Method   | Endpoint            | Description |
|----------|---------------------|-------------|
| `POST`   | `/documents`        | Insert one or many documents |
| `GET`    | `/documents/<id>`   | Fetch a document (`?include_vector=true` for the raw vector) |
| `DELETE` | `/documents/<id>`   | Delete a document |
| `POST`   | `/search`           | Single hybrid search |
| `POST`   | `/search/batch`     | Batch search |
| `GET`    | `/health`           | Document / index-node counts |
| `GET`    | `/stats`            | Full statistics |

Production:

```bash
gunicorn -w 4 -b 0.0.0.0:8000 "akshamdb.server:create_app()"
```

---

## How it works

| Layer | Implementation |
|-------|----------------|
| **Storage** | LSM-style store: write-ahead log (fsync) → in-memory memtable → SSTables with Bloom filters and background compaction. |
| **Index** | C++ HNSW graph (Malkov & Yashunin, 2018) via pybind11, GIL released during search. |
| **Query** | Metadata filter → ANN retrieval → LSM fetch → hybrid cosine + Okapi BM25 scoring, then optional MMR / calibration / reranking. |
| **Scale (opt)** | Product Quantization for memory compression; IVF coarse index for very large corpora. |

On open, the index is loaded from disk when it matches the stored documents,
and only rebuilt (fully or incrementally) when it does not — so startup stays
cheap as the database grows.

---

## Logging

`akshamdb` uses the standard `logging` module and is silent until you add a
handler:

```python
import logging
logging.basicConfig(level=logging.INFO)
logging.getLogger("akshamdb").setLevel(logging.WARNING)   # or quiet it
```

---

## License

Apache-2.0
