Metadata-Version: 2.4
Name: langchain-gridgain
Version: 3.0.0
Summary: Use GridGain as a Vector Store, Document Loader, LLM Cache, Key-Value Store and Chat Memory within LangChain. For LangGraph persistence, see langgraph-checkpoint-gridgain.
Author-email: Manini Puranik <manini.puranik@gridgain.com>, Aditi Sharma <aditi.sharma@gridgain.com>
Project-URL: Homepage, https://www.gridgain.com/
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Database
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: langchain-core<2,>=1.4.7
Requires-Dist: pygridgain<2.0,>=1.6.0
Requires-Dist: pyignite<0.7,>=0.6.1
Provides-Extra: test
Requires-Dist: langchain-classic<2,>=1.0.8; extra == "test"
Requires-Dist: langchain-tests<2,>=1.1.9; extra == "test"
Requires-Dist: pytest>=9; extra == "test"
Requires-Dist: pytest-asyncio; extra == "test"
Requires-Dist: pytest-cov; extra == "test"
Requires-Dist: pytest-mock; extra == "test"
Requires-Dist: pytest-socket; extra == "test"
Requires-Dist: pytest-timeout; extra == "test"
Provides-Extra: integration
Requires-Dist: testcontainers>=4.15; extra == "integration"
Provides-Extra: bench
Requires-Dist: pytest-benchmark>=5.2; extra == "bench"
Requires-Dist: langgraph-checkpoint<5,>=4.1.1; extra == "bench"
Requires-Dist: hypothesis>=6.163; extra == "bench"
Requires-Dist: langgraph-checkpoint-postgres<4,>=3.1; extra == "bench"
Requires-Dist: psycopg[binary,pool]>=3.2; extra == "bench"
Requires-Dist: testcontainers>=4.15; extra == "bench"
Provides-Extra: lint
Requires-Dist: ruff>=0.16; extra == "lint"
Requires-Dist: mypy>=2.3; extra == "lint"
Requires-Dist: codespell>=2.4.3; extra == "lint"

# langchain-gridgain

langchain-gridgain is a Python library that provides seamless integration between GridGain/Apache Ignite and LangChain. This library offers a set of storage adapters that allow LangChain components to efficiently use GridGain as a backend for various data storage needs.

## Table of Contents
1. [Features](#features)
2. [Prerequisites](#prerequisites)
3. [Installation](#installation)
   - [LangGraph users](#langgraph-users)
4. [GridGain Setup](#gridgain-setup)
   - [Connecting to GridGain](#1-connecting-to-gridgain)
5. [Detailed Component Explanations](#detailed-component-explanations)
   - [GridGainStore](#1-gridgainstore)
   - [GridGainDocumentLoader](#2-gridgaindocumentloader)
   - [GridGainChatMessageHistory](#3-gridgainchatmessagehistory)
   - [GridGainCache](#4-gridgaincache)
   - [GridGainSemanticCache](#5-gridgainsemanticcache)
   - [GridGainVectorStore](#6-gridgainvectorstore)
   - [GridGainByteStore](#7-gridgainbytestore)
   - [LangGraph persistence moved out in 3.0.0](#langgraph-persistence-moved-out-in-300)
6. [Entry Expiry (TTL)](#entry-expiry-ttl)
7. [Upgrading to 3.0.0](#upgrading-to-300)
8. [Upgrading to 2.0.0](#upgrading-to-200)
9. [Performance](#performance)
10. [Documentation](#documentation)
11. [Example](#example)

## Features

This library implements key LangChain and LangGraph interfaces for GridGain:

1. **GridGainStore**: A key-value store implementation.
2. **GridGainDocumentLoader**: A document loader for retrieving documents from GridGain caches.
3. **GridGainChatMessageHistory**: A chat message history store using GridGain.
4. **GridGainCache**: A caching mechanism for Language Models using GridGain.
5. **GridGainSemanticCache**: A semantic caching mechanism for Language Models using GridGain.
6. **GridGainVectorStore**: A vector store implementation using GridGain for storing and querying embeddings.
7. **GridGainByteStore**: A binary key-value store, e.g. for caching embeddings via `CacheBackedEmbeddings`.
8. **GridGainCheckpointSaver**: A LangGraph checkpointer for agent state persistence (resume, human-in-the-loop, time-travel).
9. **GridGainMemoryStore**: A LangGraph store for cross-thread agent memory, with native TTL.
10. **GridGainNodeCache**: A LangGraph node cache — a node with a `CachePolicy` runs once per distinct input.


## Prerequisites

1. Python 3.10 or above (3.11, 3.12 and 3.13 are tested)
    * You can use `pyenv` to manage multiple Python versions (optional):
        1. Install `pyenv`: `brew install pyenv` (or your system's package manager)
        2. Create and activate the environment: 
            ```bash
            pyenv virtualenv 3.11.7 langchain-env
            source $HOME/.pyenv/versions/langchain-env/bin/activate 
            ```
    * Alternatively, ensure supported Python version is installed directly.

2. A running GridGain node, at least 8.9.17 ([release notes](https://www.gridgain.com/docs/latest/release-notes/8.9.17/release-notes_8.9.17)). Which edition you need depends on what you use:
   - **Community Edition is enough** for `GridGainCheckpointSaver`, `GridGainMemoryStore`, `GridGainNodeCache`, `GridGainCache`, `GridGainStore`, `GridGainByteStore`, `GridGainChatMessageHistory` and `GridGainDocumentLoader` — they are key-value and SQL only.
   - **Enterprise or Ultimate with a vector-search licence** is required for `GridGainVectorStore` and `GridGainSemanticCache`.

## Installation

Install the package using pip:

```bash
pip install langchain-gridgain
```

### LangGraph users

The LangGraph checkpointer and store are a separate package:

```bash
pip install langgraph-checkpoint-gridgain
```

```python
from langgraph.cache.gridgain import GridGainNodeCache
from langgraph.checkpoint.gridgain import GridGainCheckpointSaver
from langgraph.store.gridgain import GridGainMemoryStore
```

It depends on this package, so that one install gives you both the LangGraph
persistence and the LangChain components. See
[Upgrading to 3.0.0](#upgrading-to-300) if you are moving from 2.x.

## GridGain Setup

In order to use [GridGain](https://www.gridgain.com/) powered langchain components, you need a running GridGain cluster. Vector search is required only for `GridGainVectorStore` and `GridGainSemanticCache`; every other component, `GridGainNodeCache` included, runs on Community Edition (see [Prerequisites](#prerequisites)).

### 1. Connecting to Gridgain

```python
from pygridgain import Client

def connect_to_gridgain(host: str, port: int) -> Client:
    try:
        client = Client()
        client.connect(host, port)
        print("Connected to Ignite successfully.")
        return client
    except Exception as e:
        print(f"Failed to connect to Ignite: {e}")
        raise
```

Usage:
```python
client = connect_to_gridgain("localhost", 10800)
```

## Detailed Component Explanations

### 1. GridGainStore

GridGainStore is a key-value store implementation that uses GridGain as its backend. It provides a simple and efficient way to store and retrieve data using key-value pairs.

Usage example:
```python
from langchain_gridgain.storage import GridGainStore

def initialize_keyvalue_store(client) -> GridGainStore:
    try:
        key_value_store = GridGainStore(
            cache_name="laptop_specs",
            client=client
        )
        print("GridGainStore initialized successfully.")
        return key_value_store
    except Exception as e:
        print(f"Failed to initialize GridGainStore: {e}")
        raise

# Usage
client = connect_to_ignite("localhost", 10800)
key_value_store = initialize_keyvalue_store(client)

# Store a value
key_value_store.mset([("laptop1", "16GB RAM, NVIDIA RTX 3060, Intel i7 11th Gen")])

# Retrieve a value
specs = key_value_store.mget(["laptop1"])[0]
```

### 2. GridGainDocumentLoader

GridGainDocumentLoader is designed to load documents from GridGain caches. It's particularly useful for scenarios where you need to retrieve and process large amounts of textual data stored in GridGain.

Usage example:
```python
from langchain_gridgain.document_loaders import GridGainDocumentLoader

def initialize_doc_loader(client) -> GridGainDocumentLoader:
    try:
        doc_loader = GridGainDocumentLoader(
            cache_name="review_cache",
            client=client,
            create_cache_if_not_exists=True
        )
        print("GridGainDocumentLoader initialized successfully.")
        return doc_loader
    except Exception as e:
        print(f"Failed to initialize GridGainDocumentLoader: {e}")
        raise

# Usage
client = connect_to_ignite("localhost", 10800)
doc_loader = initialize_doc_loader(client)

# Populate the cache
reviews = {
    "laptop1": "Great performance for coding and video editing. The 16GB RAM and dedicated GPU make multitasking a breeze."
}
doc_loader.populate_cache(reviews)

# Load documents
documents = doc_loader.load()
```

### 3. GridGainChatMessageHistory

GridGainChatMessageHistory provides a way to store and retrieve chat message history using GridGain. This is crucial for maintaining context in conversational AI applications.

Usage example:
```python
from langchain_gridgain.chat_message_histories import GridGainChatMessageHistory

def initialize_chathistory_store(client) -> GridGainChatMessageHistory:
    try:
        chat_history = GridGainChatMessageHistory(
            session_id="user_session",
            cache_name="chat_history",
            client=client
        )
        print("GridGainChatMessageHistory initialized successfully.")
        return chat_history
    except Exception as e:
        print(f"Failed to initialize GridGainChatMessageHistory: {e}")
        raise

# Usage
client = connect_to_ignite("localhost", 10800)
chat_history = initialize_chathistory_store(client)

# Add a message to the history
chat_history.add_user_message("Hello, I need help choosing a laptop.")

# Retrieve the conversation history
messages = chat_history.messages
```

### 4. GridGainCache

GridGainCache provides a caching mechanism for the responses received from LLMs using GridGain. This can significantly improve response times for **exact** queries by storing and retrieving pre-computed results.

Usage example:

```python
from langchain_gridgain.llm_cache import GridGainCache

def initialize_llm_cache(client)-> GridGainCache:
    try:
        llm_cache = GridGainCache(
            cache_name="llm_cache",
            client=client
        )
        logger.info("GridGainCache initialized successfully.")
        return llm_cache
    except Exception as e:
        logger.error(f"Failed to initialize GridGainCache: {e}")
        raise
```

### 5. GridGainSemanticCache

GridGainSemanticCache provides a semantic caching mechanism for the responses received from LLMs using GridGain. This can significantly improve response times for **similar** queries by storing and retrieving pre-computed results.

Usage example:

```python
from langchain_gridgain.llm_cache import GridGainCache
from langchain_gridgain.llm_cache import GridGainSemanticCache


def initialize_semantic_llm_cache(client, embedding)-> GridGainSemanticCache:
    try:
        llm_cache = GridGainCache(
            cache_name="llm_cache",
            client=client
        )
        semantic_cache = GridGainSemanticCache(
            llm_cache=llm_cache,
            cache_name="semantic_llm_cache",
            client=client,
            embedding=embedding,
            similarity_threshold=0.85
        )
        logger.info("GridGainSemanticCache initialized successfully.")
        return semantic_cache
    except Exception as e:
        logger.error(f"Failed to initialize GridGainSemanticCache: {e}")
        raise

### 6. GridGainVectorStore

GridGainVectorStore is a vector store implementation using GridGain for storing and querying embeddings. It allows efficient similarity search operations on high-dimensional vector data, and implements the standard LangChain `VectorStore` surface — it passes the `langchain-tests` conformance suite.

Usage example:
```python
from langchain_gridgain.vectorstores import GridGainVectorStore

vector_store = GridGainVectorStore(
    cache_name="tech_reviews",
    embedding=embedding_model,
    client=client,
)

texts = [
    "The latest MacBook Pro offers exceptional performance for video editing.",
    "ASUS ROG Zephyrus G14 provides a balance of portability and gaming performance.",
]

# Ids are optional: pass the standard `ids` parameter, put an "id" in metadata,
# or let the store generate them. Metadata is stored and returned untouched.
ids = vector_store.add_texts(
    texts,
    metadatas=[{"category": "laptop"}, {"category": "laptop"}],
    ids=["review-1", "review-2"],
)

# Similarity search, with or without scores
docs = vector_store.similarity_search("What's a good laptop for video editing?", k=2)
scored = vector_store.similarity_search_with_score("video editing", k=2)

# Maximal Marginal Relevance — trade relevance against diversity
diverse = vector_store.max_marginal_relevance_search(
    "laptops", k=2, fetch_k=10, lambda_mult=0.5
)

# Fetch and delete by id
vector_store.get_by_ids(["review-1"])
vector_store.delete(["review-1"])   # delete everything with delete()
```

Notes:

- **Scores are cosine distances** (`0.0` identical, larger is less similar). The vector query returns matches without their similarity values, so the distance is recomputed from each match's stored vector — the index ranks by cosine, so this is consistent with the server's own ordering. `similarity_search_with_relevance_scores` and the `similarity_score_threshold` retriever therefore work as usual; note that relevance is `1 - distance`, so genuinely opposed vectors score below `0` and LangChain warns about it.
- **Metadata filtering is not supported.** The vector query has no metadata predicate, so a `filter` could only be applied after the server had already chosen the top *k* — silently returning fewer results than asked for. A non-empty `filter` therefore raises `NotImplementedError` rather than being ignored. Partition the data (one store per tenant or collection) if you need it.
- `score_threshold` *is* applied server-side, on the engine's own similarity scale.
- **Async methods come from `VectorStore`'s defaults**, which run the sync implementation in a thread pool. That is safe, but unlike `GridGainCheckpointSaver` and `GridGainMemoryStore` it is not true non-blocking I/O.
- Deleting a document also removes it from the vector index, since the index is derived from the cache rows.
- This is the one component that **requires** a vector-enabled GridGain build and a vector-search license.

### 7. GridGainByteStore

GridGainByteStore is a binary key-value store (`BaseStore[str, bytes]`) backed by GridGain. Its main use is caching computed embeddings with LangChain's `CacheBackedEmbeddings`, so each text is embedded only once.

Usage example:
```python
from langchain_classic.embeddings import CacheBackedEmbeddings
from langchain_gridgain.storage import GridGainByteStore

byte_store = GridGainByteStore(
    cache_name="embeddings_cache",
    client=client
)

cached_embedder = CacheBackedEmbeddings.from_bytes_store(
    underlying_embeddings,
    byte_store,
    namespace="my-embedding-model",
    key_encoder="sha256",  # the default is SHA-1, which is not collision-resistant
)

# First call computes and caches; repeated calls hit GridGain.
vectors = cached_embedder.embed_documents(["Hello world"])
```

### LangGraph persistence moved out in 3.0.0

`GridGainCheckpointSaver`, `GridGainMemoryStore` and `GridGainNodeCache` are no
longer part of this package. They live in
[`langgraph-checkpoint-gridgain`](https://pypi.org/project/langgraph-checkpoint-gridgain/),
which every other LangGraph backend's naming follows:

```bash
pip install langgraph-checkpoint-gridgain
```

```python
from langgraph.cache.gridgain import GridGainNodeCache
from langgraph.checkpoint.gridgain import GridGainCheckpointSaver
from langgraph.store.gridgain import GridGainMemoryStore
```

That package depends on this one, so installing it gives you both halves. See
[its README](https://pypi.org/project/langgraph-checkpoint-gridgain/) for the
checkpointer and store documentation, and
[Upgrading to 3.0.0](#upgrading-to-300) below for what to change.

## Entry Expiry (TTL)

`GridGainCache`, `GridGainSemanticCache`, `GridGainVectorStore`, `GridGainStore` and `GridGainByteStore` accept an optional `ttl` argument (seconds or a `datetime.timedelta`). Entries written through the component expire that long after creation or update; `None` (the default) keeps entries forever.

```python
from datetime import timedelta
from langchain_gridgain.llm_cache import GridGainCache, GridGainSemanticCache

llm_cache = GridGainCache(cache_name="llm_cache", client=client, ttl=timedelta(hours=1))

semantic_cache = GridGainSemanticCache(
    llm_cache=llm_cache,          # give it the same ttl so both sides expire in lockstep
    cache_name="semantic_llm_cache",
    client=client,
    embedding=embedding,
    ttl=timedelta(hours=1),
)
```

For the semantic cache, `ttl` applies to its vector entries; pass the same `ttl` to the wrapped `GridGainCache` so the exact-match entries expire in lockstep.

## Upgrading to 3.0.0

**Breaking: the LangGraph checkpointer and store moved to their own package.**

`langchain-gridgain` 3.0.0 contains the LangChain components only. If you use
`GridGainCheckpointSaver` or `GridGainMemoryStore`:

```bash
pip install langgraph-checkpoint-gridgain
```

```diff
-from langchain_gridgain import GridGainCheckpointSaver, GridGainMemoryStore
+from langgraph.cache.gridgain import GridGainNodeCache
+from langgraph.checkpoint.gridgain import GridGainCheckpointSaver
+from langgraph.store.gridgain import GridGainMemoryStore
```

Nothing else changes: the classes, their behavior and their stored data are the
same, and the new package depends on this one, so the LangChain components stay
available beside them.

Why: every other LangGraph persistence backend ships its checkpointer in a
`langgraph-checkpoint-<vendor>` distribution — postgres, sqlite, redis and
mongodb all do — and their `langchain-<vendor>` packages contain no LangGraph
code. Keeping both surfaces in one package made GridGain's checkpointer hard to
find and mixed two APIs that the ecosystem keeps apart.

If you only use the LangChain components, 3.0.0 also drops the
`langgraph-checkpoint` dependency from your install.

## Upgrading to 2.0.0

Coming from any 1.0.x release. A version-by-version summary lives in
[`CHANGELOG.md`](CHANGELOG.md); this section covers what you have to *do*.

**The dependency floor moved (breaking).** 1.0.x pinned the `langchain`
umbrella exactly (`langchain == 0.3.21`, plus `langchain-community~=0.3.20` and
`pygridgain == 1.5.0`). 2.0.0 depends on `langchain-core >= 1.4.7, < 2`
instead, drops `langchain-community` entirely, and needs `pygridgain >= 1.6`.
The old pin held LangGraph two major lines behind, which is untenable for a
package whose purpose is LangGraph integration. Because the requirement is a
hard floor, pip will simply keep resolving you to 1.0.3 until your environment
is on `langchain-core` 1.x.

**Cache keys changed (breaking).** `GridGainCache` entries are now keyed by *prompt and LLM* (previously the LLM was ignored, so the same prompt sent to different models collided on one entry). After upgrading, entries written by 1.0.x are unreachable under the new keys: the cache starts cold, and — since 1.0.x had no TTL — the old entries never expire on their own. Run `clear()` once per `GridGainCache`/`GridGainSemanticCache` after upgrading to reclaim that space:

```python
llm_cache.clear()        # exact-match cache
semantic_cache.clear()   # vector entries + its exact-match cache
```

**Embedding-cache keys changed if you followed the old `CacheBackedEmbeddings` example (breaking).** That example previously relied on LangChain's default `key_encoder`, which is SHA-1; it now passes `key_encoder="sha256"`. The key encoder *is* the cache key, so entries written under the old default become unreachable and, with no TTL on the example's `GridGainByteStore`, never expire. Either keep the old behaviour explicitly (`key_encoder="sha1"`) or clear the byte store once after upgrading — it is a `ByteStore`, so there is no `clear()`; use `byte_store.mdelete(list(byte_store.yield_keys()))`. Nothing is lost either way — the entries are recomputable — but the first run after the switch re-embeds everything.

**Semantic cache error handling (now documented).** `GridGainSemanticCache.lookup()` and `update()` degrade gracefully: a backend failure is logged and treated as a cache miss / skipped write, so a chain keeps working through GridGain hiccups and a computed generation is never lost to a cache-write error. `clear()` raises on failure. The exact-match `GridGainCache` propagates errors on all operations, as before.

**The checkpointer's schema changed, and migrates itself.** `GridGainCheckpointSaver` used to inline the serialized checkpoint into `lg_checkpoints`. It now keeps that payload in a separate `lg_checkpoint_blobs` table, because Ignite materializes whole rows during an ordered index scan — so a payload column in the table we `ORDER BY` made "fetch the latest checkpoint" cost time proportional to the *whole thread's* size. On a thread of 1000 64 KB checkpoints that was 8x slower than it needed to be.

Nothing is required of you in the ordinary case: the first saver constructed with a **sync** client against an old table copies the payloads across and drops the legacy columns, once, logging at INFO. It is idempotent and safe to run concurrently. `asetup()` does not migrate, so a saver given only an `aio_client` needs one sync construction against the same cluster to convert a legacy table. This only arises for deployments that ran the checkpointer from source before 2.0.0 — no published version has the pre-split schema. If the column drop fails (an index on it, say) you get a WARNING with the exact `ALTER TABLE` to run — reads are correct either way, they just stay slow until the columns are gone. The cost of the split is a second round trip per `put`; see [`benchmarks/RESULTS.md`](benchmarks/RESULTS.md) for what that trades against.

**`GridGainMemoryStore` creates an index.** A `(prefix, itemKey)` index is created on the store cache at construction (`CREATE INDEX IF NOT EXISTS`, so existing caches pick it up too). It backs the namespace scope, `list_namespaces`, and the `ORDER BY` that makes a paged `search` a page rather than a sort of the whole subtree.

**Client support.** The vector-based components (`GridGainVectorStore`, `GridGainSemanticCache`) require a **pygridgain** client — pyignite has no vector API. The key-value components (`GridGainStore`, `GridGainByteStore`, `GridGainCache`, `GridGainChatMessageHistory`, `GridGainDocumentLoader`) work with either client.

**Thread-safety.** The sync thin client is not thread-safe on its own — one socket, no locking — so every component built on a given client shares a single re-entrant lock and serializes its access to it. Sharing a client across components and threads is therefore safe and needs no external locking; the trade-off is that operations on one client run one at a time, so shard across several clients if you need client-side parallelism. The default async methods run their sync counterparts in a thread pool (safe for the same reason, but not true non-blocking I/O); `GridGainCheckpointSaver`, `GridGainMemoryStore` and `GridGainNodeCache` do real async I/O when given an `AioClient`. Note that read-modify-write sequences across separate calls (e.g. `GridGainChatMessageHistory.add_message`, which reads the history then writes it back) are still not atomic: the lock serializes each individual operation, not a multi-call sequence.

## Performance

Benchmarks live in [`benchmarks/`](benchmarks/README.md): pytest-benchmark
microbenchmarks of the pure-Python hot paths, and a macro sweep that drives the
checkpointer, store and caches against a real node next to two baselines —
LangGraph's `InMemorySaver` (the floor) and `langgraph-checkpoint-postgres` (the
peer). Both write JSON artifacts; a nightly workflow tracks the trend.

The one number worth knowing before you deploy: the synchronous thin client is a
single socket, so components sharing one client serialize (see
**Thread-safety** above). If you need client-side parallelism, shard across
clients — the sweep reports both shapes side by side.

## Documentation

This README is the documentation for 2.x. The GridGain docs site still
describes these components under their old `langchain_community.*` import
paths, which this package has never used and which 2.0.0 cannot satisfy at all
— `langchain-community` is no longer a dependency. It is deliberately not
linked here until it is updated.

## Example

For a comprehensive, real-world example of how to use this package, please refer to the following GitHub repository:

[GG Langchain Demo](https://github.com/GridGain-Demos/gg8_langchain_demo)

gg8_langchain_demo is a demonstration project that showcases the integration of GridGain/Apache Ignite with LangChain, using the custom langchain-gridgain package. This project provides examples of how to use GridGain as a backend for various LangChain components, focusing on a laptop recommendation system.
