Metadata-Version: 2.4
Name: sigil-telemetry
Version: 0.6.1
Summary: Plug-and-play telemetry for AI agents — records every LLM call, through any SDK or plain HTTP, and exports traces via OpenTelemetry.
Author: Zurain Khan
License: MIT
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Requires-Dist: opentelemetry-api>=1.20.0
Requires-Dist: opentelemetry-sdk>=1.20.0
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.20.0
Requires-Dist: opentelemetry-semantic-conventions>=0.41b0
Provides-Extra: anthropic
Requires-Dist: opentelemetry-instrumentation-anthropic>=0.30.0; extra == "anthropic"
Provides-Extra: openai
Requires-Dist: opentelemetry-instrumentation-openai>=0.30.0; extra == "openai"
Provides-Extra: langchain
Requires-Dist: opentelemetry-instrumentation-langchain>=0.30.0; extra == "langchain"
Provides-Extra: crewai
Requires-Dist: opentelemetry-instrumentation-crewai>=0.30.0; extra == "crewai"
Provides-Extra: llamaindex
Requires-Dist: opentelemetry-instrumentation-llamaindex>=0.30.0; extra == "llamaindex"
Provides-Extra: vertexai
Requires-Dist: opentelemetry-instrumentation-vertexai>=0.30.0; extra == "vertexai"
Provides-Extra: mistral
Requires-Dist: opentelemetry-instrumentation-mistralai>=0.30.0; extra == "mistral"
Provides-Extra: bedrock
Requires-Dist: opentelemetry-instrumentation-bedrock>=0.30.0; extra == "bedrock"
Provides-Extra: litellm
Requires-Dist: openinference-instrumentation-litellm>=0.1.0; extra == "litellm"
Provides-Extra: cohere
Requires-Dist: opentelemetry-instrumentation-cohere>=0.30.0; extra == "cohere"
Provides-Extra: fastapi
Requires-Dist: opentelemetry-instrumentation-fastapi>=0.41b0; extra == "fastapi"
Provides-Extra: flask
Requires-Dist: opentelemetry-instrumentation-flask>=0.41b0; extra == "flask"
Provides-Extra: django
Requires-Dist: opentelemetry-instrumentation-django>=0.41b0; extra == "django"
Provides-Extra: celery
Requires-Dist: opentelemetry-instrumentation-celery>=0.41b0; extra == "celery"
Provides-Extra: all
Requires-Dist: opentelemetry-instrumentation-anthropic>=0.30.0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-openai>=0.30.0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-langchain>=0.30.0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-crewai>=0.30.0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-llamaindex>=0.30.0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-vertexai>=0.30.0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-mistralai>=0.30.0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-bedrock>=0.30.0; extra == "all"
Requires-Dist: openinference-instrumentation-litellm>=0.1.0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-cohere>=0.30.0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-fastapi>=0.41b0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-flask>=0.41b0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-django>=0.41b0; extra == "all"
Requires-Dist: opentelemetry-instrumentation-celery>=0.41b0; extra == "all"

# sigil-telemetry

Plug-and-play telemetry for AI agents. Install it, call `init()`, and every LLM call your agent makes is automatically tracked.

## Quick Start

```bash
pip install sigil-telemetry[all]
```

```python
from sigil_telemetry import init
init()
```

```bash
# Set your agent's ID and collector endpoint
SIGIL_AGENT_ID=your-agent-slug
SIGIL_COLLECTOR_URL=https://your-collector-endpoint/
```

That's it. Every LLM API call is now captured — tokens, model, latency, errors — and sent to your collector.

---

## What's New in v0.6.1

- **List attributes now reach Sijil.** The Azure Monitor exporter silently drops list-valued attributes, so `sigil.input.files`, `sigil.handoff.from` and instrumentors' lists such as `gen_ai.response.finish_reasons` never arrived. Every list attribute is now exported as a JSON array string, e.g. `sigil.input.files = ["image · png · 88 KB"]`.
- **Prompts of calls an instrumentor skips.** When an SDK instrumentor records a call but leaves out its input (the OpenAI instrumentor does for vision calls with images), the prompt the transport read is added — images and files as their type only (`{"type": "image_url"}`), never their data. What the instrumentor recorded is never replaced, and nothing is added when content capture is off (`SIGIL_CAPTURE_CONTENT`, `TRACELOOP_TRACE_CONTENT`, `OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT` or `OPENINFERENCE_HIDE_INPUTS`).
- **Launcher scripts aren't the agent.** An agent started from a console script (`uvicorn`, `gunicorn` in the venv's `bin/`) no longer gets `__main__.<module>` as its `sigil.agent.site` when its own code is installed as a package.

## What's New in v0.6.0

- **Every LLM call is captured, however it is made.** Before 0.6.0 only calls through an instrumented SDK were seen — an agent calling a model with plain `httpx.post(...)` sent nothing. `init()` now also watches the HTTP clients underneath (`httpx` and its fork `httpx2`, which the OpenAI SDK uses from 3.x, `requests`/`urllib3`, `boto3`, `aiohttp`, `urllib`) plus `websockets` (OpenAI Realtime) and `grpc` (older Gemini/Vertex SDKs). A call is recognised from its host, its route, or its request/response shape — Anthropic, OpenAI (Chat and Responses), Azure OpenAI and Foundry, Gemini/Vertex, Bedrock, Ollama, Cohere and any OpenAI-compatible server — streamed (SSE, NDJSON, AWS event-stream) or not, compressed or not.
- **No double counting.** When an SDK instrumentor is already recording a call, transport capture records no span of its own; SDK spans keep their richer detail (tool calls, agent names) and get the agent attributes below.
- **Prompts and answers recorded** on every LLM span (`gen_ai.input.messages`, `gen_ai.output.messages`, `gen_ai.system_instructions`, tool calls and results included). What Sigil receives is kept as valid JSON under 8000 characters, below the Application Insights property limit of 8192 — the latest turn is protected, older turns are shortened or dropped first. Every other long string on every span, including the instrumentors' own prompt copies (`gen_ai.prompt.N.content`, `llm.input_messages.N…`, `input.value`), is cut to 8000 too, ending in `…[truncated]`. (If the app also exports to its own backend, that export gets the content unshortened, as the SDK instrumentors always sent it.) Turn off with `SIGIL_CAPTURE_CONTENT=false`.
- **One attribute vocabulary.** Spans from older OpenLLMetry versions (`gen_ai.system`, `gen_ai.usage.prompt_tokens`, `gen_ai.prompt.N.*`) and from OpenInference (`llm.*`, e.g. LiteLLM) are exported with the current `gen_ai.*` names too, so a different instrumentor version can't silently change Sigil's numbers.
- **Works alongside the app's own tracing.** If the Azure Monitor distro, Logfire or anything else already set a TracerProvider, Sigil's exporter is added to it instead of being silently dropped.
- **Processes too.** `ProcessPoolExecutor` and `multiprocessing.Process` work keeps the request's trace and user, and is flushed when it ends. Celery tasks are traced from the code that queued them (`[celery]` extra).
- **Agents told apart, from what they do.** Every LLM span says which of the agent's functions made it (`sigil.agent.site`, e.g. `app.services.llm._llm_extract ← …`). Calls from one site are compared, and what stays the same across them is the agent's **fixed prompt**: after 5 calls with at least 3 different inputs (and 2 different users, when calls carry a user) the span gets `sigil.prompt.status = split` and `sigil.prompt.id` — one id per AI task, so the same instructions called from two places in the code are one agent, and different instructions are different agents. The fixed prompt is sent once per process per version (`sigil.prompt.fixed.0`, `.1`, …, scrubbed of emails, phone numbers, IBANs, card and national id numbers), and from then on each run's `gen_ai.input.messages` carries only its input, with `{"type": "fixed"}` where the fixed prompt was.
- **Non-text input described** in `sigil.input.files`: `pdf · Renewal_2026.pdf · 2.1 MB`, `image · png · 88 KB`, `audio · wav · 388 KB · 12.4 s` (seconds exact for WAV and Realtime audio, `~` estimated for MP3). Files uploaded to start a request (FastAPI `UploadFile`, Flask `request.files`, Django `request.FILES`) are described on the request span — name, type and size only; this follows `SIGIL_CAPTURE_USER`. `sigil.input.chars` is the whole prompt's length before any cutting.
- **Handoffs between agents.** `sigil.handoff.from` lists the earlier spans (`<trace id>-<span id>`) whose answer this call's input carries (at least 30% of it) — one agent's output becoming another's input. `sigil.handoff.in` / `sigil.handoff.out` / `sigil.handoff.in_max` are hashed fingerprints of the input and answer, so your backend can find handoffs across processes and hours (a person reviewing in between). Ids, sites, file descriptions and fingerprints are sent even with `SIGIL_CAPTURE_CONTENT=false`; no text is.

## What's New in v0.3.0

- **Automatic user identity capture** — Extracts the calling user from Azure AD / SSO headers or JWT Bearer tokens on every request. No code changes needed — works out of the box with `init()`. Disable with `SIGIL_CAPTURE_USER=false`.
- **Worker agent support** — Set `SIGIL_USER_ID` env var for scheduled jobs and queue workers that don't have HTTP requests.
- **Background work stays in its run** (0.5.2) — LLM calls made on a `ThreadPoolExecutor` or `Thread` after the response has been sent keep the trace and user of the request that started them. Before 0.5.2 each became its own orphan trace with no user.

## What's New in v0.2.0

- **Web framework auto-instrumentation** — FastAPI, Flask, and Django are auto-detected and instrumented. All LLM calls within one HTTP request share a single `trace_id` (`operation_Id`), so you can count agent "runs" with `COUNT(DISTINCT trace_id)`.
- **Noise span filtering** — `http send` / `http send body` spans from web frameworks are silently dropped before they leave the process. They never reach your collector.
- **Health check exclusion** — Routes like `/health`, `/healthz`, `/ready`, `/alive`, `/ping` are excluded from tracing entirely.
- **Graceful shutdown** — `atexit` handler flushes all pending spans when the process exits, so you never lose the last batch.
- **Lighter install** — Removed unnecessary dependencies from the core install.

---

## Full Example

Here's a real agent that summarizes documents using Claude:

### 1. The Agent Code

```python
# document_summarizer.py
import anthropic
from sigil_telemetry import init

# Initialize telemetry — call this ONCE at startup
init()

# Your normal agent code — no changes needed
client = anthropic.Anthropic()
response = client.messages.create(
    model="claude-sonnet-4-20250514",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Summarize this document: ..."}
    ]
)
print(response.content[0].text)
```

### 2. What Gets Captured (Per LLM Call)

Every time `client.messages.create()` runs, a **span** is automatically created with:

| Field | Example Value | Description |
|-------|--------------|-------------|
| `operation_Id` | `a1b2c3d4e5f6...` | Trace ID — groups all LLM calls in a single agent run |
| `sigil.agent.id` | `doc-summarizer` | Which agent made the call |
| `sigil.agent.version` | `0.1.0` | Agent version |
| `sigil.agent.frameworks` | `Anthropic,FastAPI` | Which SDKs and frameworks were detected |
| `sigil.user.id` | `jane.doe@example.com` | Who triggered this run (auto-captured) |
| `gen_ai.system` | `anthropic` | LLM provider |
| `gen_ai.request.model` | `claude-sonnet-4-20250514` | Model used |
| `gen_ai.usage.input_tokens` | `1250` | Tokens sent |
| `gen_ai.usage.output_tokens` | `340` | Tokens received |
| `duration` | `2.3s` | How long the call took |
| `status` | `OK` or `ERROR` | Whether the call succeeded |
| `sigil.environment` | `production` | Environment |
| `sigil.agent.division` | `Sales` | Business division (if set) |
| `sigil.agent.risk_classification` | `low` | Risk level (if set) |

If the agent makes **multiple LLM calls** in one run (e.g., calls Claude then GPT-4), all calls share the same `operation_Id` so you can see the full trace.

### 3. How Trace Grouping Works

**API agents (FastAPI/Flask/Django):** The web framework instrumentor creates a root span per HTTP request. All LLM calls within that request automatically become child spans sharing the same `trace_id`. You don't need to do anything — `init()` handles it.

**Worker agents (scheduled jobs, listeners):** Wrap your `main()` function in a custom span. All LLM calls within one job share the same `trace_id`.

In both cases: `COUNT(DISTINCT trace_id)` = number of agent runs.

### 4. Where the Data Goes

```
Agent makes LLM call
        │
        ▼
sigil-telemetry auto-captures it as an OpenTelemetry span
(noise spans like "http send" are filtered out here)
        │
        ▼
Span is batched and sent via OTLP to:
  → Your configured collector endpoint
        │
        ▼
Collector forwards to:
  → Your observability backend (Jaeger, Zipkin, Datadog, etc.)
```

---

## User Identity Capture

v0.3.0 automatically captures who triggered each agent run. No code changes needed.

**How it works:**

| Auth Method | How Identity Is Captured |
|-------------|------------------------|
| Azure AD / SSO (Easy Auth) | Reads `X-MS-CLIENT-PRINCIPAL-NAME` header automatically |
| Azure AD / SSO (in-app JWT) | Decodes Bearer token, extracts `preferred_username` → `upn` → `email` → `name` → `oid` |
| Worker agents (no HTTP) | Reads `SIGIL_USER_ID` env var |

The captured user ID is attached as `sigil.user.id` on every request span. All LLM calls within that request inherit the user context through the shared trace.

**Disable user capture:**

```bash
SIGIL_CAPTURE_USER=false
```

Or in code:

```python
init(SigilConfig(capture_user=False))
```

---

## Supported SDKs

Use `[all]` to install everything. Only the SDKs your agent actually uses get activated.

| SDK | Install Extra | What It Covers |
|-----|--------------|----------------|
| Anthropic | `[anthropic]` | Anthropic API |
| OpenAI | `[openai]` | OpenAI API (including compatible endpoints) |
| LangChain | `[langchain]` | LangChain, LangGraph, any LangChain-wrapped model |
| CrewAI | `[crewai]` | CrewAI multi-agent framework |
| LlamaIndex | `[llamaindex]` | LlamaIndex agents and pipelines |
| Vertex AI | `[vertexai]` | Google Vertex AI, Gemini models |
| Mistral AI | `[mistral]` | Mistral API |
| AWS Bedrock | `[bedrock]` | Claude, Llama, Titan via AWS |
| LiteLLM | `[litellm]` | Unified proxy across 100+ LLM providers |

**Anything else** — raw HTTP, an SDK without an instrumentor, a self-hosted or gateway endpoint — is recorded by transport capture with no extra install: model, tokens, prompt, answer, latency and errors (no SDK-level detail such as tool calls). Those spans carry `sigil.capture.source = transport`.

Not covered: agents that are not written in Python, low-code platforms, models running inside the process (`transformers`, `llama-cpp`), Batch APIs (usage arrives hours later in a file), and gateways whose responses don't carry token usage (the call is recorded without tokens when the host is a known model host).

## Web Framework Auto-Instrumentation

These are included in `[all]` and auto-detected by `init()`:

| Framework | Install Extra | What It Does |
|-----------|--------------|--------------|
| FastAPI | `[fastapi]` | Creates root span per HTTP request — all LLM calls in that request share one trace_id |
| Flask | `[flask]` | Same trace grouping for Flask apps |
| Django | `[django]` | Same trace grouping for Django apps |
| Celery | `[celery]` | One trace per task, linked to the code that queued it |

Health check routes (`/health`, `/healthz`, `/ready`, `/alive`, `/ping`, `/startup`, `/liveness`, `/readiness`) are automatically excluded from tracing.

## Configuration

| Env Variable | Default | Description |
|-------------|---------|-------------|
| `SIGIL_AGENT_ID` | — | **Required.** Your agent's unique ID |
| `SIGIL_AGENT_VERSION` | `0.1.0` | Track deployments |
| `SIGIL_COLLECTOR_URL` | — | **Required.** Your collector endpoint URL |
| `SIGIL_ENVIRONMENT` | `production` | `production`, `staging`, `development` |
| `SIGIL_CONSOLE_EXPORT` | `false` | Print spans to console for debugging |
| `SIGIL_DIVISION` | — | Business division (e.g., `Sales`, `Engineering`) |
| `SIGIL_RISK_CLASSIFICATION` | — | Agent risk level (`low`, `medium`, `high`) |
| `SIGIL_HOURS_SAVED` | — | Estimated hours saved per run |
| `SIGIL_CAPTURE_USER` | `true` | Set `false` to disable automatic user identity capture |
| `SIGIL_USER_ID` | — | Static user ID for worker agents with no HTTP context |
| `SIGIL_PROPAGATE_CONTEXT` | `true` | Set `false` to stop background threads and processes inheriting the request's trace and user |
| `SIGIL_CAPTURE_CONTENT` | `true` | Set `false` to stop recording prompts and answers |
| `SIGIL_CAPTURE_TRANSPORT` | `true` | Set `false` to record only calls made through an instrumented SDK |
| `SIGIL_MAX_CONTENT_CHARS` | `8000` | Longest string attribute sent. Can be lowered, never raised past 8000 (Application Insights cuts or drops a property over 8192) |

Or pass config in code:

```python
from sigil_telemetry import init, SigilConfig

init(SigilConfig(
    agent_id="my-agent",
    environment="development",
    console_export=True,
    capture_user=True
))
```

## Custom Spans

Track things beyond LLM calls (document parsing, tool use, etc.):

```python
from sigil_telemetry import get_tracer, record_error

tracer = get_tracer()

with tracer.start_as_current_span("parse-contract") as span:
    span.set_attribute("document.pages", 42)
    try:
        result = parse_pdf(file)
    except Exception as e:
        record_error(span, e)
        raise
```
