Metadata-Version: 2.5
Name: raindrop-langchain
Version: 0.0.13
Summary: Raindrop integration for LangChain
Project-URL: Homepage, https://raindrop.ai
Project-URL: Documentation, https://docs.raindrop.ai
Author-email: Raindrop AI <sdk@raindrop.ai>
License-Expression: MIT
License-File: LICENSE
Classifier: Typing :: Typed
Requires-Python: >=3.10
Requires-Dist: langchain-core>=0.3.0
Requires-Dist: raindrop-ai>=0.0.70
Provides-Extra: dev
Requires-Dist: pytest>=8.0; extra == 'dev'
Description-Content-Type: text/markdown

# raindrop-langchain

Raindrop integration for LangChain (Python). Automatically captures LLM calls, tool usage, chains, and retrievers via LangChain's callback system.

## Installation

```bash
pip install raindrop-langchain langchain-core
```

## Quick Start

```python
from raindrop_langchain import RaindropLangchain
from langchain_openai import ChatOpenAI
from langchain_core.messages import HumanMessage

raindrop = RaindropLangchain(
    api_key="rk_...",
    user_id="user-123",
)

model = ChatOpenAI(model="gpt-4o")

result = model.invoke(
    [HumanMessage(content="Hello!")],
    config={"callbacks": [raindrop.handler]},
)

raindrop.flush()
```

### Factory Function (alternative)

```python
from raindrop_langchain import create_raindrop_langchain

raindrop = create_raindrop_langchain(api_key="rk_...", user_id="user-123")
model = ChatOpenAI(model="gpt-4o")
result = model.invoke("Hello!", config={"callbacks": [raindrop.handler]})
raindrop.flush()
```

## Projects

Route events to a specific [project](https://docs.raindrop.ai/platform/projects) by passing its slug as `project_id`:

```python
raindrop = RaindropLangchain(
    api_key="rk_...",
    project_id="support-prod",
)
```

`project_id` sets the `X-Raindrop-Project-Id` header on every event. Omit it (or pass `"default"`) to use your org's default **Production** project, which is the existing behavior. The same option is accepted by the `create_raindrop_langchain(...)` factory. Invalid slugs are ignored with a warning and no header is sent.

## What Gets Captured

- **LLM calls** — model name, input, output, token usage, finish reason
- **Tool calls** — tool name, input arguments, output, duration (via `interaction.track_tool()` spans)
- **Chains** — execution tracking
- **Retrievers** — query and document count
- **Errors** — error type and message captured in event properties
- **Extended token categories** — cached tokens (`ai.usage.cached_tokens`) and reasoning tokens (`ai.usage.thoughts_tokens`) when available from the provider (e.g. OpenAI)
- **Finish reason** — captured as `ai.finish_reason` in event properties (e.g. `"stop"`, `"length"`)

## Debug Mode

Enable verbose logging with `debug=True`:

```python
raindrop = RaindropLangchain(
    api_key="rk_...",
    debug=True,
)
```

## Identify Users

Associate events with a user after initialization:

```python
raindrop.identify("user-123", {"name": "Alice", "plan": "pro"})
```

## Track Signals

Send feedback, edits, or custom signals:

```python
raindrop.track_signal(
    event_id="evt-abc",
    name="thumbs_up",
    signal_type="feedback",
    sentiment="POSITIVE",
)
```

## Flushing and Shutdown

```python
raindrop.flush()     # flush pending data
raindrop.shutdown()  # flush + release resources
```

## API Reference

### `RaindropLangchain`

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `api_key` | `Optional[str]` | `None` | Raindrop API key. If `None`, telemetry is disabled |
| `user_id` | `Optional[str]` | `None` | Associate all events with a user |
| `convo_id` | `Optional[str]` | `None` | Group events into a conversation |
| `project_id` | `Optional[str]` | `None` | Route events to a specific [project](https://docs.raindrop.ai/platform/projects) (slug); omit for the default **Production** project |
| `trace_chains` | `bool` | `True` | Track chain execution |
| `trace_retrievers` | `bool` | `True` | Track retriever calls |
| `filter_langgraph_internals` | `bool` | `True` | Filter LangGraph-internal chain events and deduplicate LLM callbacks |
| `event_name` | `str` | `"ai_generation"` | Event name applied to every event (the `event` field in the dashboard) |
| `tracing_enabled` | `bool` | `True` | Enable distributed tracing |
| `bypass_otel_for_tools` | `bool` | `True` | Bypass OTEL for tool spans |
| `debug` | `bool` | `False` | Enable debug logging |

#### Methods

| Method | Description |
|--------|-------------|
| `handler` | Property — the LangChain callback handler to pass into `config={"callbacks": [...]}` |
| `flush()` | Flush all pending events to the Raindrop API |
| `shutdown()` | Flush remaining events and release resources |
| `identify(user_id, traits)` | Identify a user with optional traits |
| `track_signal(event_id, name, ...)` | Track a signal event |

## Async Support

The callback handler inherits from LangChain's `AsyncCallbackHandler` and works with both synchronous and asynchronous LangChain invocations.

```python
result = await model.ainvoke(
    [HumanMessage(content="Hello!")],
    config={"callbacks": [raindrop.handler]},
)
```

## LangGraph Support

Works with LangGraph out of the box. The handler automatically filters LangGraph-internal chain events and deduplicates LLM callbacks. Pass the handler to the model inside your LLM node — **not** to `graph.invoke()`. See `examples/langchain-langgraph-python-basic/` for a full example.

## LangSmith Coexistence

Raindrop and LangSmith can run simultaneously. Set `LANGSMITH_TRACING=false` to disable LangSmith if you only want Raindrop.

## Payload size bounds

Payloads the handler serializes itself — multi-modal chat content lists and
agent-action tool inputs — are bounded to 1,000,000 characters with a
`...[truncated by raindrop]` marker. The bound is enforced *during*
serialization (cost proportional to the cap, not the payload), so a multi-MB
content list (e.g. base64 image parts) can't stall your event loop inside a
synchronous callback. Plain-string prompts and tool outputs are capped by
the Raindrop SDK's own per-field limit (`max_text_field_chars`, raindrop-ai
>= 0.0.51).

## Tool catalog capture (`ai.prompt.tools`)

Every model span records the tools LangChain handed the provider for that call as `ai.prompt.tools`: an OTel string array, one JSON document per tool with `name`, `description` and the full JSON Schema as `inputSchema` (`{"type": "function", ...}`), or `{"type": "provider-defined", ...}` for provider-owned tools. The capture is **exact, per call**: `RaindropCallbackHandler.on_chat_model_start` / `on_llm_start` receive `invocation_params["tools"]` (or the legacy `functions`), which is the list `bind_tools()` / `bind(tools=...)` put on the request. The handler binds that list through the core's `prompt_tools_scope` mechanism for the duration of the model call and releases it in `on_llm_end` / `on_llm_error`, so the model span the OpenLLMetry LangChain instrumentor (or the OpenAI instrumentor under it) starts inside the call carries the attribute. A call whose `invocation_params` has no `tools` or `functions` records `[]`. Sync, async (`ainvoke`, the handler runs inline in the caller's task) and streaming calls are covered because the span starts inside the call.

Callback and instrumentor agree: the instrumentor writes the same list to `gen_ai.tool.definitions` (OpenLLMetry 0.62 emits that attribute, not `llm.request.functions.*`; `raindrop-ai` 0.0.70/0.0.71 do not rebuild `ai.prompt.tools` from it and a core patch is in flight, but the wrapper binds the list itself so capture does not depend on it). The callback's `invocation_params` is the source of truth; the test suite asserts both carry the same tools.

There is no model span when the instrumentor is off, so there is nothing to carry the attribute. `RaindropLangchain(...)` builds its own `Raindrop` client with `auto_instrument=False`; pass `client=Raindrop(..., tracing_enabled=True, instruments={Instruments.LANGCHAIN})` (or `auto_instrument=True`) to get model spans.

Override:

```python
# Default for every model span this handler sees
rd = RaindropLangchain(api_key="...", user_id="u1", tools=[{"type": "function", "name": "get_weather", "inputSchema": {...}}])
handler = RaindropCallbackHandler(user_id="u1", tools=[...])          # same option on the raw handler

# Per call, through LangChain config metadata (stripped from the event properties)
chain.invoke(messages, config={"callbacks": [handler], "metadata": {"raindrop_tools": [...]}})
```

The per-call value replaces the constructor default; the constructor default replaces the inferred list; `tools=[]` at either level records an empty catalog. An enclosing `raindrop.prompt_tools(...)` / `begin(..., tools=...)` scope is also treated as an override and is never shadowed by the inferred list.

Semantics shared with every Raindrop SDK:

- **Attribute absent**: the wrapper could not see the tool list (or there is no model span, see below).
- **`[]`**: the model had no tools.
- **Override replaces, never merges.** A `tools=` you pass wins over anything inferred; `tools=[]` records an empty catalog.
- **Content gate.** Tool definitions are prompt content. With `TRACELOOP_TRACE_CONTENT=false` nothing is recorded, overrides included.
- Malformed entries are skipped with a debug log; nothing here raises into your application.

Accepted shapes for `tools=`: canonical `{"type": "function", "name", "description", "inputSchema"}`, OpenAI `{"type": "function", "function": {...}}`, Anthropic `{"name", "description", "input_schema"}`, and `{"type": "provider-defined", "name", ...}`. The core normalises them; the JSON Schema is kept exactly as given.

Needs `raindrop-ai>=0.0.70`; on an older core the wrapper logs at debug and records nothing.

## Known Limitations

- **Multi-LLM chain data**: In ReAct loops with multiple child LLMs, only the last child's data survives (Python SDK uses one-shot `track_ai` vs TS's accumulative `EventShipper.patch`).
- **Error-path input loss**: On LLM errors, the input captured during `on_llm_start` is not forwarded to the finalized event.

## Application Git metadata

`RaindropLangchain(...)` and `create_raindrop_langchain(...)` accept the keyword-only `app_git` option. It defaults to `True`: explicit Raindrop Git environment or deployment context is applied immediately, and the base SDK may perform one bounded background local-Git lookup from the process working directory. Event capture, flush, and shutdown never wait for that lookup. Pass `False` to disable enrichment, or pass an `AppGitOptions` mapping with `commit_sha`, `commit_dirty`, `branch`, `source_directory`, `detect_branch`, and/or `auto_detect`. Automatic branch discovery remains opt-in through `detect_branch=True` (or `RAINDROP_GIT_DETECT_BRANCH=true`).

For an ordinary in-process application, the process working directory is treated as the application-under-test checkout. A remote, coding, workflow, or observer process must not rely on its own checkout: pass `app_git=False`, provide explicit revision values, or set `source_directory` to the actual application checkout. Canonical per-operation properties remain authoritative. When supplying `client=`, configure `app_git` while constructing that `Raindrop` client; the supplied client is authoritative and the wrapper's `app_git` argument does not reconfigure it.

Release order is deliberate: first publish the base SDK feature, then publish the wrapper feature release with its minimum dependency coordinated to that base release. The existing `raindrop-ai` lower bound remains compatible, but application Git metadata is unavailable on an older core and must not be claimed complete until the base is upgraded. Until coordination assigns a released version, the wrapper checks for an explicit base `app_git` parameter and omits the option when unsupported. Explicit non-default configuration is debug-logged and omitted. Unsupported `app_git` is determined by signature inspection before construction, not by retrying initialization after a `TypeError`; Git configuration adds no initialization attempts and does not change any existing framework-specific initialization fallback. `RaindropCallbackHandler(...)` is unchanged and does not accept this option.

## Testing

```bash
cd packages/langchain-python
pip install -e .
python -m pytest tests/ -v   # unit tests (no external services)
```

End-to-end behavior is verified by the cross-SDK **conformance harness**. This
package ships a thin conformance driver at
[`conformance/driver.py`](conformance/driver.py) that maps the shared scenario
corpus onto the wrapper's public API; known gaps are tracked as ticket-linked
entries in [`conformance/failures.txt`](conformance/failures.txt). The **fault
lane** runs on every PR touching `packages/*-python/**`
(`.github/workflows/conformance-wrappers-python.yml`) against a local capture
server; the **prod lane** verifies delivery by reading back through the public
[Query API](https://docs.raindrop.ai/api-reference/overview). The harness is
pinned by commit SHA (`HARNESS_REF`). See the harness docs:
[HOW-IT-WORKS](https://github.com/invisible-tools/raindrop-sdk-harness/blob/main/docs/HOW-IT-WORKS.md) ·
[AGENTS](https://github.com/invisible-tools/raindrop-sdk-harness/blob/main/AGENTS.md) ·
[README](https://github.com/invisible-tools/raindrop-sdk-harness/blob/main/README.md).
