Metadata-Version: 2.4
Name: llmpivot
Version: 0.3.0
Summary: Versioned prompts and Agent Skills with progressive context loading for production LLM apps
Author-email: Sanath Goutham <sanathgoutham@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/Lancer59/llmpivot
Project-URL: Repository, https://github.com/Lancer59/llmpivot
Keywords: prompts,agent-skills,llm,progressive-context,fastapi,versioning,mongodb,rbac,instruction-studio
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Framework :: FastAPI
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: fastapi>=0.110.0
Requires-Dist: httpx>=0.27.0
Requires-Dist: aiosqlite>=0.20.0
Requires-Dist: python-multipart>=0.0.9
Provides-Extra: server
Requires-Dist: uvicorn>=0.29.0; extra == "server"
Provides-Extra: mongo
Requires-Dist: motor>=3.3.0; extra == "mongo"
Requires-Dist: pymongo>=4.6.0; extra == "mongo"
Provides-Extra: security
Requires-Dist: argon2-cffi>=23.1.0; extra == "security"
Provides-Extra: all
Requires-Dist: uvicorn>=0.29.0; extra == "all"
Requires-Dist: motor>=3.3.0; extra == "all"
Requires-Dist: pymongo>=4.6.0; extra == "all"
Requires-Dist: argon2-cffi>=23.1.0; extra == "all"
Dynamic: license-file

# llmpivot

**Versioned prompts and Agent Skills for production LLM applications — with a built-in AI assistant.**

Change a prompt in the web UI — see it reflected in your running app in seconds. No redeployment. No code changes. And when you need help, the **Assistant** is already watching the screen.

---

## What it does

llmpivot sits alongside your FastAPI app. You mount it at a path (e.g. `/prompts`) and it gives you:

- A web UI to create, edit, version, compare, test, and activate prompts and Agent Skills
- Standard `SKILL.md` bundles with metadata, Markdown references, and ZIP import/export
- Progressive skill loading: expose descriptions first, then load one skill or reference only when needed
- An in-memory cache so your app reads prompts at zero latency
- Full version history with one-click rollback
- RBAC authentication (admin / editor / viewer)
- A **prompt hierarchy** — organise prompts as agent → tool → utility trees
- Per-prompt **metadata** (purpose, type, owner, sensitivity, call location)
- Auto-generated **changelogs** on every version save
- An **Application Context** document so the Assistant understands your whole system
- A **floating Assistant widget** that watches what you're doing and surfaces observations, alerts, and suggestions without being asked
- A **fallback snapshot** (`prompts_fallback.json`) written on every activation — if the DB goes down your app keeps serving the last known good prompts
- Usage logging, A/B testing, audit log, import/export, and a real `/healthz` endpoint
- Lightweight prompt token estimates in Instruction Studio, with no tokenizer dependency

Your app code just does:

```python
meta = await aget_prompt_with_meta("my_prompt")
# use meta["content"] as your LLM system prompt
```

Instruction Studio shows a rough token estimate for each active prompt and saved
version. You can also call `estimate_tokens(text)` from Python. It uses a
lightweight four-characters-per-token heuristic; the model tokenizer remains the
source of truth for billing and context limits.

---

## Installation

```bash
# Standard (SQLite, recommended for most setups)
pip install llmpivot

# With strong password hashing (recommended for production)
pip install "llmpivot[security]"

# With MongoDB backend
pip install "llmpivot[mongo]"

# Everything
pip install "llmpivot[all]"
```

Or from the repo:

```bash
pip install -r requirements.txt
pip install -e .
```

---

## Quick start

```python
import os
from fastapi import FastAPI, Request
from fastapi.responses import JSONResponse
from llmpivot import LLMAssetManager, PromptNotFoundError, aget_prompt_with_meta, log_prompt_usage

manager = LLMAssetManager(
    db_path="prompts.db",
    cache_ttl=5,
    auth_mode="rbac",
    secret_key=os.environ["LLMPIVOT_SECRET"],
    cookie_secure=True,
    bootstrap_admin=True,
    bootstrap_password=os.environ.get("LLMPIVOT_PASSWORD", "changeme"),

    # LLM — enables AI suggestions, A/B testing, and the Assistant widget
    llm_url="https://api.openai.com/v1/chat/completions",
    llm_api_key=os.environ["OPENAI_API_KEY"],
    llm_model="gpt-4o",

    # Assistant widget
    pivot_enabled=True,
    pivot_proactive=True,
    auto_changelog=True,

    # Fallback snapshot
    fallback_snapshot=True,
)

app = FastAPI()

@app.exception_handler(PromptNotFoundError)
async def _not_found(request: Request, exc: PromptNotFoundError):
    return JSONResponse(status_code=404, content={"error": "prompt_not_found", "detail": str(exc)})

app.mount("/prompts", manager.mount_ui())

@app.get("/run")
async def run(text: str):
    meta = await aget_prompt_with_meta("my_prompt")
    output = call_your_llm(meta["content"], text)
    log_prompt_usage("my_prompt", meta["version_id"], text, output)
    return {"output": output}
```

Run:

```bash
uvicorn example_app:app --reload
```

Open **http://localhost:8000/prompts/list** — log in with `admin` / `changeme`.

---

## Azure OpenAI

Pass the Azure endpoint directly — llmpivot detects Azure URLs automatically, uses the correct `api-key` header, and selects `max_completion_tokens` vs `max_tokens` based on the model generation:

```python
manager = LLMAssetManager(
    db_path="prompts.db",
    llm_url="https://myresource.openai.azure.com/openai/deployments/gpt-4o/chat/completions?api-version=2024-08-01-preview",
    llm_api_key=os.environ["AZURE_OPENAI_API_KEY"],
    llm_model="gpt-4o",           # used for model-generation detection
    llm_api_type="azure",          # explicit; or omit and let auto-detection handle it
    pivot_enabled=True,
)
```

Or let `example_app.py` handle everything from your `.env`:

```
AZURE_OPENAI_ENDPOINT=https://myresource.openai.azure.com/
AZURE_OPENAI_API_KEY=...
AZURE_OPENAI_DEPLOYMENT_NAME=gpt-4o
AZURE_OPENAI_API_VERSION=2024-08-01-preview
```

---

## The Assistant widget

When `pivot_enabled=True`, a floating **Assistant** pill appears in the bottom-right corner of every dashboard page. It has two modes:

**Proactive** (default) — opens a page and the Assistant immediately analyses it:
- Flags prompts with no active version (red warning)
- Alerts when a parent prompt changed but its children haven't been reviewed (amber)
- Suggests adding metadata or changelog entries (blue)
- Shows version count, last editor, and sensitivity level as context (grey)

**Chat** — type a question in the input at the bottom of the expanded panel:
- "What does this prompt do and where is it called?"
- "What would break if I changed the tone here?"
- "Make this more specific about the output format" → returns a diff preview, never saves without confirmation
- "Show me all prompts under the support agent"

The widget state (open/closed) is persisted per browser in `localStorage`. No page refresh needed.

## Agent Skills

Skills are managed as versioned LLM instruction assets alongside prompts. Create them in the UI, save new versions, compare and evaluate versions side by side, activate a release, and import or export portable skill ZIPs. A skill bundle contains a standard `SKILL.md` and may include Markdown files under `references/`. Bundled scripts and binary assets are not executed or loaded.

Application code can progressively load a short catalog, then one selected skill, then a named reference:

```python
from llmpivot import alist_skills, aget_skill, aget_skill_reference

catalog = await alist_skills()  # names and descriptions only
skill = await aget_skill("release-review")  # active SKILL.md + version metadata
policy = await aget_skill_reference("release-review", "references/policy.md")
```

`LLMAssetManager.skill_tool_schemas()` and `await manager.call_skill_tool(name, arguments)` provide a provider-neutral function-tool bridge for agent frameworks. Pivot uses progressive loading by default. Set `skill_loading_mode="eager"` to include all active skill instructions and references in each Pivot chat request.

---

## Prompt hierarchy

Prompts can have parent-child relationships. This lets you model real multi-agent systems:

```
customer_support_agent    [agent]
├── classify_intent       [tool]
├── draft_response        [tool]
└── escalation_check      [tool]
```

Set hierarchy in **Metadata** (`/prompts/metadata/{name}`). The tree view is at `/prompts/tree`. When you edit a parent prompt, the Assistant automatically warns you about child prompts that may need review.

---

## Fallback snapshot

`prompts_fallback.json` lives next to your database and contains all currently active prompt contents:

```json
{
  "my_prompt": "You are a helpful assistant...",
  "classify_intent": "Classify the user intent into..."
}
```

Written atomically on every activation and on startup. Use `aget_prompt_with_fallback` instead of `aget_prompt` for zero-downtime resilience:

```python
from llmpivot import aget_prompt_with_fallback

# Tries live DB/cache first; serves from snapshot if DB is unreachable
content = await aget_prompt_with_fallback("my_prompt")
```

---

## Configuration reference

All parameters passed to `LLMAssetManager()`.

`PromptManager` remains an import-compatible alias. New applications should use `LLMAssetManager`.

### Storage

| Parameter | Type | Default | Description |
|---|---|---|---|
| `db_path` | `str` | `"prompts.db"` | SQLite file path |
| `storage_type` | `str` | `"sqlite"` | `"sqlite"` or `"mongodb"` |
| `mongo_uri` | `str` | `None` | MongoDB connection URI |
| `mongo_db_name` | `str` | `"llmpivot"` | MongoDB database name |
| `tenant_id` | `str` | `"default"` | Namespace for multi-tenant isolation |
| `cache_ttl` | `int` | `5` | Seconds before cache re-fetches from DB |

### Auth

| Parameter | Type | Default | Description |
|---|---|---|---|
| `auth_mode` | `str` | `"disabled"` | `"disabled"`, `"protected"`, or `"rbac"` |
| `secret_key` | `str` | **required for rbac** | HMAC key for session cookie signing |
| `cookie_secure` | `bool` | `True` | HTTPS-only session cookie. Set `False` for local HTTP dev |
| `bootstrap_admin` | `bool` | `False` | Create `admin` user on first startup |
| `bootstrap_password` | `str` | `"admin"` | Password for the bootstrapped admin |

### Logging

| Parameter | Type | Default | Description |
|---|---|---|---|
| `log_sample_rate` | `float` | `1.0` | Fraction of calls to log (0.0–1.0) |

### LLM

| Parameter | Type | Default | Description |
|---|---|---|---|
| `llm_url` | `str` | `None` | Chat completions endpoint URL |
| `llm_api_key` | `str` | `None` | API key |
| `llm_model` | `str` | `"gpt-3.5-turbo"` | Model name (also used for token-field auto-detection) |
| `llm_api_type` | `str` | auto | `"openai"` or `"azure"`. Auto-detected from URL if omitted |
| `llm_max_tokens` | `int` | `None` | Set for GPT-4 and earlier. Auto-selected if both are `None` |
| `llm_max_completion_tokens` | `int` | `None` | Set for GPT-5 / o1 / o3+. Auto-selected if both are `None` |
| `llm_temperature` | `float` | `None` | Omitted from request if `None` (new-gen models reject it) |
| `llm_top_p` | `float` | `None` | Omitted from request if `None` |
| `llm_timeout` | `float` | `30.0` | Request timeout in seconds |
| `llm_extra_params` | `dict` | `None` | Extra body params, e.g. `{"response_format": {"type": "json_object"}}` |
| `llm_suggester_prompt` | `str` | `None` | Override the system prompt for the AI suggest feature |

### Assistant

| Parameter | Type | Default | Description |
|---|---|---|---|
| `pivot_enabled` | `bool` | `False` | Master on/off switch. `False` = no widget, no LLM calls |
| `pivot_proactive` | `bool` | `True` | Auto-analyse current page. `False` = chat-only mode |
| `auto_changelog` | `bool` | `True` | Auto-generate changelog entry on every version save (requires LLM) |
| `pivot_model` | `str` | `None` | Override model for the Assistant specifically |
| `skill_loading_mode` | `str` | `"progressive"` | Pivot context policy: load individual skills on demand, or eagerly include all active skills |

### Fallback snapshot

| Parameter | Type | Default | Description |
|---|---|---|---|
| `fallback_snapshot` | `bool` | `True` | Write `prompts_fallback.json` on activation and startup |
| `fallback_path` | `str` | `None` | Override file path. Default: same directory as `db_path` |

---

## Auth modes

### `"disabled"` (default)
No authentication. Use only behind a VPN or on localhost.

### `"protected"`
Single admin password required for all write operations.

```python
LLMAssetManager(protected_mode=True, admin_password="your-password")
```

### `"rbac"` (recommended for production)
Full multi-user RBAC with login/logout, session cookies, and role enforcement.

```python
LLMAssetManager(
    auth_mode="rbac",
    secret_key=os.environ["LLMPIVOT_SECRET"],
    bootstrap_admin=True,
    bootstrap_password="first-login-password",
)
```

**Roles:**

| Role | Access |
|---|---|
| `admin` | Everything: user management, delete prompts, import/export, edit, activate |
| `editor` | Create and edit versions, activate versions, import, edit metadata and context |
| `viewer` | Read-only: view prompts, versions, diffs, logs, export, browse hierarchy |

---

## Web UI routes

| Route | Auth | Description |
|---|---|---|
| `/prompts/list` | viewer | All prompts, active versions, last editor |
| `/prompts/tree` | viewer | Prompt hierarchy tree view |
| `/prompts/edit/__new__` | editor | Create a new prompt |
| `/prompts/edit/{name}` | editor | Add a new version (hierarchy warning shown if children exist) |
| `/prompts/detail/{name}` | viewer | Version history, activate/rollback |
| `/prompts/metadata/{name}` | editor | Edit prompt metadata and hierarchy |
| `/prompts/changelog/{name}` | viewer | Per-version changelog feed |
| `/prompts/diff/{name}` | viewer | Side-by-side diff between any two versions |
| `/prompts/test/{name}` | viewer | A/B test two versions with live LLM calls |
| `/prompts/context` | editor | Application Context document (read/write) |
| `/prompts/health` | viewer | Prompt health dashboard |
| `/prompts/search` | viewer | Semantic search across all prompts |
| `/prompts/logs` | viewer | Usage log viewer |
| `/prompts/import` | editor | Bulk import from JSON |
| `/prompts/export` | viewer | Download active prompts as JSON |
| `/prompts/users` | admin | Create users, assign roles |
| `/prompts/login` | — | Login page |
| `/prompts/logout` | — | Clears session cookie |
| `/prompts/healthz` | — | Liveness probe (DB + worker health) |
| `/prompts/pivot/observe` | viewer | SSE stream: proactive observations for current page |
| `/prompts/pivot/chat` | viewer | Chunked HTTP: Assistant chat reply stream |
| `/prompts/skills` | viewer | Active Agent Skill catalog |
| `/prompts/skills/edit/{name}` | editor | Create a new skill version |
| `/prompts/skills/detail/{name}` | viewer | Skill history and activation |
| `/prompts/skills/diff/{name}` | viewer | Compare two SKILL.md versions |
| `/prompts/skills/test/{name}` | viewer | Evaluate two skill versions against the same task |
| `/prompts/skills/import` and `/prompts/skills/export/{name}` | editor / viewer | Import or export portable skill ZIPs |

---

## Python API

### `aget_prompt(name) → str`
Returns the active prompt content. Preferred in async routes.

### `aget_prompt_with_meta(name) → dict`
Returns `{"content": str, "version_id": int}`. Use this when logging.

### `aget_prompt_with_fallback(name) → str`
Tries live cache/DB first. Falls back to `prompts_fallback.json` if unavailable. Raises `PromptNotFoundError` only if absent from both.

### `get_prompt(name) → str`
Sync wrapper. Use in plain scripts or non-async contexts.

### `log_prompt_usage(name, version_id, input_text, output_text)`
Fire-and-forget usage logging. Non-blocking. When a workflow must wait for queued
entries to appear in the Logs view, call `await manager.usage_logger.flush()`.

### Agent Skill loading

- `alist_skills()` returns active skill names and descriptions without loading their instructions.
- `aget_skill(name)` returns one active `SKILL.md` and its version metadata.
- `aget_skill_reference(name, path)` loads one Markdown reference under `references/` on demand.
- `LLMAssetManager.skill_tool_schemas()` and `call_skill_tool()` expose a provider-neutral progressive function-tool flow for application agents.

---

## Security notes

- **`secret_key`** must be set explicitly when `auth_mode="rbac"`. The default raises `ValueError`. Generate: `python -c "import secrets; print(secrets.token_hex(32))"`
- **`cookie_secure`** defaults to `True` — HTTPS-only. Set `False` only for local HTTP dev.
- **CSRF protection** on all state-mutating POST requests. Tokens are session-bound.
- **Login rate limiting** — 10 failed attempts per IP per 5 minutes → 15-minute lockout.
- **Password hashing** — Argon2id → bcrypt → PBKDF2 (auto-selected, backward compatible).
- **Audit log** — every create, activate, delete, import, and user-create action is recorded.

---

## Production checklist

- [ ] `secret_key` loaded from environment variable or secrets manager
- [ ] `cookie_secure=True` — app served over HTTPS
- [ ] `bootstrap_password` changed after first login
- [ ] `argon2-cffi` installed: `pip install argon2-cffi`
- [ ] `auth_mode="rbac"`
- [ ] `/prompts/healthz` wired to load balancer / k8s liveness probe
- [ ] `cache_ttl` ≤ 5s for multi-worker deployments
- [ ] `log_sample_rate` reduced if log volume is high (e.g. `0.1`)
- [ ] `prompts.db` on a persistent volume with scheduled backups
- [ ] `prompts_fallback.json` included in deployment for offline resilience

---

## License

[MIT](LICENSE) © Sanath Goutham
