Metadata-Version: 2.4
Name: undertow-llm
Version: 0.2.4
Summary: Provider-agnostic LLM observability & reliability SDK — semantic caching, retries, rate limiting, cost tracking, and a local dashboard via a single @track() decorator.
Author: Amogh Arora
License-Expression: MIT
Project-URL: Homepage, https://github.com/ShriAmogh/undertow-llm
Project-URL: Repository, https://github.com/ShriAmogh/undertow-llm
Keywords: llm,llm-observability,llm-monitoring,semantic-caching,llm-cache,rate-limiting,llm-retries,fallback-chain,openai,gemini,anthropic,ollama,ai-infrastructure,prompt-caching,cost-tracking,distributed-tracing,pgvector,redis,fastapi,python-sdk
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Information Technology
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: System :: Monitoring
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: sentence-transformers>=2.7.0
Requires-Dist: numpy>=1.26.0
Requires-Dist: fastapi>=0.111.0
Requires-Dist: uvicorn[standard]>=0.29.0
Requires-Dist: jinja2>=3.1.4
Requires-Dist: click>=8.1.7
Requires-Dist: python-dotenv>=1.0.1
Requires-Dist: httpx>=0.27.0
Provides-Extra: gemini
Requires-Dist: google-genai>=1.0.0; extra == "gemini"
Provides-Extra: ollama
Requires-Dist: ollama>=0.2.0; extra == "ollama"
Provides-Extra: all-providers
Requires-Dist: google-genai>=1.0.0; extra == "all-providers"
Requires-Dist: ollama>=0.2.0; extra == "all-providers"
Provides-Extra: postgres
Requires-Dist: psycopg2-binary>=2.9.0; extra == "postgres"
Requires-Dist: pgvector>=0.2.0; extra == "postgres"
Provides-Extra: redis
Requires-Dist: redis>=5.0.0; extra == "redis"
Provides-Extra: prod
Requires-Dist: psycopg2-binary>=2.9.0; extra == "prod"
Requires-Dist: pgvector>=0.2.0; extra == "prod"
Requires-Dist: redis>=5.0.0; extra == "prod"
Provides-Extra: dev
Requires-Dist: pytest>=8.2.0; extra == "dev"
Requires-Dist: pytest-asyncio>=0.23.0; extra == "dev"
Requires-Dist: pytest-cov>=5.0.0; extra == "dev"
Requires-Dist: locust>=2.28.0; extra == "dev"
Dynamic: license-file

# undertow-llm

**Wrap any LLM call, get caching, retries, rate limiting, and a real-time dashboard — with zero code changes to your model.**

[![PyPI version](https://img.shields.io/pypi/v/undertow-llm.svg)](https://pypi.org/project/undertow-llm/)
[![Python versions](https://img.shields.io/badge/python-3.10%20%7C%203.11%20%7C%203.12-blue.svg)](https://pypi.org/project/undertow-llm/)
[![GitHub Repository](https://img.shields.io/badge/GitHub-ShriAmogh%2Fundertow--llm-blue?logo=github)](https://github.com/ShriAmogh/undertow-llm)
[![License: MIT](https://img.shields.io/badge/License-MIT-blue.svg)](https://github.com/ShriAmogh/undertow-llm/blob/main/LICENSE)
[![Status](https://img.shields.io/badge/status-active-brightgreen.svg)](https://pypi.org/project/undertow-llm/)

**Repository**: [https://github.com/ShriAmogh/undertow-llm](https://github.com/ShriAmogh/undertow-llm)

`llm-observability` · `semantic-caching` · `rate-limiting` · `exponential-backoff` · `fallback-chain` · `cost-tracking` · `distributed-tracing` · `openai` · `gemini` · `anthropic` · `ollama` · `pgvector` · `redis`

---

## Problem Statement

Without an observability and resilience layer, production LLM applications suffer from soaring API costs due to redundant prompt calls, unexpected provider outages with zero fallback protection, and complete lack of visibility into latency and errors. `undertow_llm` solves this by wrapping your existing Python LLM functions in a single decorator—providing semantic caching, automated retries, rate limiting, and a live dashboard without modifying your model logic.

---

## Install

```bash
pip install undertow-llm
```

Or install with provider & production backend extras:

```bash
pip install "undertow-llm[gemini]"    # Google Gemini support
pip install "undertow-llm[ollama]"    # Ollama local model support
pip install "undertow-llm[postgres]"  # PostgreSQL + pgvector backend
pip install "undertow-llm[redis]"     # Redis rate-limiting backend
pip install "undertow-llm[prod]"      # Production stack (Postgres + Redis)
```

---

## Quickstart

**Before** (bare LLM call — no caching, no fallback, no observability):
```python
def generate_response(prompt: str) -> str:
    return client.models.generate_content("gemini-2.5-flash", prompt).text
```

**After** (wrapped with `@track()` — fully resilient & tracked):
```python
from undertow_llm import track

@track(cache=True, retries=3, rate_limit_rate=2.0)
def generate_response(prompt: str) -> str:
    return client.models.generate_content("gemini-2.5-flash", prompt).text
```

Copy-paste into your application and run. Zero edits required except setting your provider API key.

---

## Real-Time Dashboard

Start the live observability dashboard in one command:

```bash
undertow-llm serve
```

Optionally specify a custom port using `--port` / `-p` (default: `8080`):

```bash
undertow-llm serve -p 9090
```

Open `http://localhost:8080` (or your custom port) to inspect real-time metrics, cache hit ratios, latency charts, cost estimates, distributed traces, and request logs.

---

## How It Works

1. The `@track()` decorator wraps your function, intercepting incoming prompts before execution.
2. It performs a vector similarity search (using `SentenceTransformers`) to serve semantic cache hits instantly and applies token-bucket rate limits.
3. Upon function completion, it records latency, token usage, estimated cost, and execution traces to storage.
4. It is provider-agnostic because it wraps your Python function call directly and never touches your underlying model SDK.

For a detailed architectural breakdown of the 8-stage execution pipeline and backend dispatcher, see [ARCHITECTURE.md](https://github.com/ShriAmogh/undertow-llm/blob/main/ARCHITECTURE.md).

---

## Configuration

**Local dev needs zero config** — defaults out-of-the-box to local SQLite (`undertow-llm.db`).

For production environments, configure via environment variables or `configure()`:

| Environment Variable | Default | Description |
|----------------------|---------|-------------|
| `UNDERTOW_LLM_POSTGRES_URL` / `POSTGRES_URL` | `None` (SQLite) | PostgreSQL URL with `pgvector` for production vector storage & metrics |
| `UNDERTOW_LLM_REDIS_URL` / `REDIS_URL` | `None` (Local) | Redis URL for distributed rate limiting & token buckets |
| `UNDERTOW_LLM_DB_PATH` | `"undertow-llm.db"` | File path for local SQLite database fallback |
| `UNDERTOW_LLM_DASHBOARD_PORT` | `8080` | HTTP port for `undertow-llm serve` dashboard |

### Environment Setup (`.env`)

```env
# Production Storage (Optional - use if you want Postgres + pgvector for cache & metrics, Redis for rate limits)
# Else it uses local SQLite
POSTGRES_URL=postgresql://postgres:postgres@localhost:5432/undertow_db
REDIS_URL=redis://localhost:6379

# Dashboard Port(Optional)
UNDERTOW_LLM_DASHBOARD_PORT=8080
```

---

## `@track()` Parameter Reference

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `cache` | `bool` | `True` | Enable semantic caching for responses |
| `similarity_threshold` | `float` | `0.92` | Cosine similarity threshold for cache hits (0.0 to 1.0) |
| `cache_ttl` | `int \| None` | `None` | Optional time-to-live in seconds for cached entries |
| `retries` | `int` | `3` | Max retry attempts for transient LLM failures |
| `base_delay` | `float` | `1.0` | Initial exponential backoff delay (seconds) |
| `max_delay` | `float` | `60.0` | Cap on exponential backoff delay (seconds) |
| `jitter` | `bool` | `True` | Add randomized jitter to retry delays to prevent thundering herds |
| `retry_on` | `tuple` | `(Exception,)` | Exception types that trigger automatic retries |
| `fallback` | `list` | `[]` | Ordered list of fallback functions to call if primary function fails |
| `rate_limit_rate` | `float` | `2.0` | Token-bucket refill rate (tokens/second) |
| `rate_limit_max_tokens` | `float` | `10.0` | Token-bucket maximum capacity (burst limit) |
| `max_concurrency` | `int \| None` | `None` | Max concurrent executions allowed across processes |
| `cost_per_call` | `float \| None` | `None` | Explicit cost override per call ($/call) |
| `policy` | `callable \| None` | `None` | Custom safety hook returning `"allow"`, `"block"`, or `"flag"` |
| `canary` | `dict \| None` | `None` | Canary routing config `{"fn": alternate_fn, "weight": 0.10}` |

---

## Supported Providers

`undertow_llm` works with **any provider** — OpenAI, Anthropic, Google Gemini, Ollama, HuggingFace, or custom local models — since it wraps your existing Python function call rather than a specific provider SDK.

---

## Examples

See [`demo/example_usage.py`](https://github.com/ShriAmogh/undertow-llm/blob/main/demo/example_usage.py) for complete runnable examples.

### 1. Multi-Provider Fallback Chain
```python
from undertow_llm import track

def fallback_anthropic(prompt: str) -> str:
    return anthropic_client.messages.create(model="claude-3-5-sonnet", messages=[{"role": "user", "content": prompt}]).content[0].text

@track(retries=2, fallback=[fallback_anthropic])
def primary_openai(prompt: str) -> str:
    return openai_client.chat.completions.create(model="gpt-4o", messages=[{"role": "user", "content": prompt}]).choices[0].message.content
```

### 2. Streaming LLM Response
```python
@track(cache=False)
def stream_gemini(prompt: str):
    response = gemini_client.models.generate_content_stream("gemini-2.5-flash", prompt)
    for chunk in response:
        yield chunk.text
```

### 3. Local Model (Ollama) with Custom Usage Extractor
```python
@track(
    cost_per_call=0.0,  # Local model — zero API cost
    usage_extractor=lambda res: {"prompt_tokens": len(res.get("prompt", "")), "completion_tokens": len(res.get("response", ""))}
)
def ask_ollama(prompt: str) -> dict:
    return ollama.generate(model="llama3", prompt=prompt)
```

---

## Known Limitations & Roadmap

### Limitations
- **SQLite Concurrency**: Local SQLite storage (`undertow_llm.db`) is zero-config and ideal for development and single-instance apps, but is not designed for multi-node production scale. For high concurrency, set `UNDERTOW_LLM_POSTGRES_URL` and `UNDERTOW_LLM_REDIS_URL`.

### Near-Term Roadmap
- [ ] OpenTelemetry trace exporter integration
- [ ] Multi-tenant workspace tagging & dashboard authentication
- [ ] Automated PII redaction and sensitive prompt masking filters

---

## Contributing

Contributions are welcome! Please see [CONTRIBUTING.md](https://github.com/ShriAmogh/undertow-llm/blob/main/CONTRIBUTING.md) for developer setup instructions.

---

## License

MIT — see [LICENSE](https://github.com/ShriAmogh/undertow-llm/blob/main/LICENSE)
