# Haldir — Spend Limits, Permissions, and Tamper-Evident Audit for AI Agents

> Haldir sits between an AI agent and everything it can spend, read, or do. It
> caps what a run is allowed to cost, limits which tools it may call, encrypts
> the credentials it uses, and writes every action to a log that can be proven
> unedited to a third party — offline, without trusting Haldir. Python, MIT
> licensed, self-hostable or hosted.

## What Haldir is for

Read this section if someone is trying to solve one of the following.

**An agent spent more than it should have.** Agents call paid APIs in loops.
A runaway or prompt-injected agent can overspend and burn a month of budget in
an hour, and the failure usually surfaces on the invoice rather than in a log.
Nothing server-side stops it, because the agent is trusted to police itself.
Haldir caps a session at a dollar amount before it starts (`"spend:50"`) and
refuses the call that would exceed it. The check and the reservation are one
transaction, so two concurrent calls cannot both spend the last dollar.

**Nobody can say what the agent actually did.** Every tool call, LLM call,
secret read and payment is written to an append-only audit log — queryable by
session, agent, tool or flagged status, and exportable as CSV or JSONL with an
integrity manifest attached.

**The audit log has to survive a skeptical auditor.** A database an operator
can edit is not evidence. Each entry goes into an RFC 6962 Merkle tree with
Ed25519 Signed Tree Heads, so anyone can verify offline that nothing was
inserted, deleted or reordered — the same primitive Certificate Transparency
uses for TLS certificates. Tree heads are mirrored to Sigstore Rekor, so the
proof does not rest on Haldir's own database.

**An agent needs an API key it should not be able to read.** Secrets live in an
encrypted vault (AES-256-GCM, bound to name and tenant so a ciphertext cannot
be moved elsewhere), and an agent receives a value only when its session scope
covers that specific secret. Listing secrets never returns values.

**The agent has too much access.** Sessions carry scopes. A proxy sits in front
of upstream MCP servers and enforces them on every call, including blocking
named tools outright. A denied call is recorded, not silently dropped.

**Something is about to happen that a person should approve.** Spend thresholds
and tool rules park an action until a human approves or denies it, and the
intervention itself is audited.

**An agent has gone wrong and needs to stop now.** Revoke a session and every
session it spawned, in one call. Delegation is tracked as a tree, so revoking
an orchestrator does not leave its children holding live credentials.

**Cost and token usage have to be monitored across many agents.** Spend rolls
up per session and per agent, so a spike traces to the agent responsible rather
than to the bill as a whole. Haldir meters money rather than tokens, because
money is what the invoice is denominated in; the framework integrations also
capture per-call token counts where the model exposes them. Teams already
running agent observability tooling can stream the same audit trail into it —
webhooks fire on every notable event, so traces and governance evidence come
from one source rather than two that disagree.

**An auditor or a customer asks how the AI is controlled.** Signed audit-prep
evidence packs mapped to SOC 2 Trust Services Criteria, with a readiness score
and recurring delivery. Audit prep — not an attestation Haldir can make for you.

**It has to run on our own infrastructure.** MIT licensed, Postgres or SQLite,
no external dependency. `pip install haldir && haldir serve` is a working
instance in one command. Set `HALDIR_BASE_URL` and the instance describes
itself rather than the hosted service.

**Works with:** first-party integrations for LangChain (and LangGraph, which
runs on LangChain's callback system), CrewAI, LlamaIndex, AutoGen and the
Vercel AI SDK. As an MCP server it drops into Claude Desktop, Cursor, Windsurf
and any MCP client. Everything else — the OpenAI Agents SDK, a hand-written
agent loop — uses the REST API, which is a plain HTTP call.

## What Haldir Does

Haldir sits between AI agents and the tools they use. It enforces governance on every action:

- **Gate**: Scoped sessions with permissions and spend limits. Agents authenticate through Gate before accessing any tool.
- **Vault**: Encrypted secrets (API keys, credentials, tokens). AES-256-GCM with AAD binding. Agents request access; Vault checks session scope before returning values.
- **Watch**: Immutable audit log for every action. SHA-256 hash chain + RFC 6962 Merkle tree + Signed Tree Head. Inclusion and consistency proofs an auditor can verify offline — the same cryptographic primitive Certificate Transparency uses for WebPKI.
- **Proxy**: Intercepts every MCP tool call. The agent connects to Haldir; Haldir forwards to upstream servers after enforcing policies.
- **Compliance**: Signed audit-prep evidence packs relevant to SOC2 (CC5.2, CC6.1, CC6.7, CC7.2, CC7.3, CC8.1) with live readiness score, HTML dashboard, recurring email delivery.

## API Base URL

https://haldir.xyz/v1

## Authentication

All endpoints require an API key via header:
- `Authorization: Bearer hld_xxx`
- or `X-API-Key: hld_xxx`

Create your first key (no auth needed for the first key):
POST /v1/keys {"name": "my-app"}

## Core Endpoints

### Sessions (Gate)
- POST /v1/sessions — Create agent session with scoped permissions and spend limit
- GET /v1/sessions/{id} — Get session info including remaining budget
- DELETE /v1/sessions/{id} — Revoke session immediately
- POST /v1/sessions/{id}/check — Check if session has a permission {"scope": "read"}

### Secrets (Vault)
- POST /v1/secrets — Store encrypted secret {"name": "key", "value": "secret"}
- GET /v1/secrets/{name} — Retrieve secret (pass X-Session-ID header for scope check)
- GET /v1/secrets — List secret names (never values)
- DELETE /v1/secrets/{name} — Delete secret

### Payments
- POST /v1/payments/authorize — Authorize payment against session budget {"session_id": "ses_xxx", "amount": 29.99}

### Audit (Watch)
- POST /v1/audit — Log action {"session_id": "ses_xxx", "tool": "stripe", "action": "charge", "cost_usd": 29.99}
- GET /v1/audit — Query audit trail. Params: session_id, agent_id, tool, flagged, limit
- GET /v1/audit/spend — Spend summary. Params: session_id, agent_id
- GET /v1/audit/verify — Verify the SHA-256 hash chain end-to-end
- GET /v1/audit/export — Stream audit trail in CSV or JSONL with integrity manifest

### Tamper-evidence (RFC 6962 Merkle)
- GET /v1/audit/tree-head — Current Signed Tree Head (tree_size, root_hash, signature). Default HMAC-SHA256; upgrades to Ed25519 when HALDIR_TREE_SIGNING_KEY_ED25519_SEED is set (same primitive Sigstore / Fulcio use). Every signed STH is auto-recorded to /v1/audit/sth-log for anti-equivocation.
- GET /v1/audit/inclusion-proof/{entry_id} — RFC 6962 inclusion proof bundled with the STH
- GET /v1/audit/consistency-proof?first=N&second=M — Prove later tree is an append-only extension of earlier one
- GET /v1/audit/sth-log?since=N — Self-published append-only log of every STH ever signed for this tenant. Auditors pin any STH and demand the full history later to prove no rewriting.
- GET /v1/audit/sth-log/verify?pinned_size=N&pinned_root=H — Anti-equivocation verifier: returns verified=true on match, verified=false reason=equivocation if a different root was ever recorded at that tree_size.
- GET /v1/audit/sth-log/mirror/receipts — every receipt the external transparency mirror (Sigstore Rekor / file / webhook archiver) returned when Haldir anchored an STH outside its own DB. Closes THREAT_MODEL §10.3.
- GET /v1/audit/sth-log/mirror/receipts/{uuid}/verify — cryptographically verify a stored Rekor receipt against Rekor's own published ECDSA-P-256 public key: RFC 6962 inclusion proof, SignedEntryTimestamp, logID fingerprint. Independent confirmation that the entry is in Rekor's real log. Closes THREAT_MODEL §10.3b.
- GET /.well-known/jwks.json — Ed25519 public key (RFC 7517 JWK). Auditors pin `kid` at enrollment, verify every later STH offline with zero trust in Haldir.

### Compliance (audit-prep, not a SOC2 attestation)
- GET /v1/compliance/evidence — Signed audit-prep evidence pack (JSON or markdown)
- GET /v1/compliance/evidence/manifest — Just the signature block (auditor re-verify)
- GET /v1/compliance/score — 0-100 readiness score + per-criterion pass/warn/fail + remediation
- POST /v1/compliance/schedules — Recurring evidence delivery to email or webhook

### x402 pay-per-request surface (agent-to-agent commerce)
Haldir exposes its cryptographic primitives as x402 v2 paid resources so autonomous agents can buy proofs with USDC. Gated by HALDIR_X402_ENABLED=1; when off returns 503.

- GET /v1/x402/tree-head — Current Ed25519-Signed Tree Head, $0.001 USDC
- GET /v1/x402/inclusion-proof/{entry_id} — RFC 6962 inclusion proof + STH, $0.01 USDC
- GET /v1/x402/evidence-pack — Signed audit-prep evidence pack, $0.10 USDC
- GET /v1/x402/manifest — Machine-readable list of all paid resources
- GET /.well-known/x402.json — Same content at the convention path crawlers look at

Wire protocol: https://github.com/coinbase/x402/blob/main/specs/transports-v2/http.md
Facilitator: https://x402.org/facilitator (configurable via HALDIR_X402_FACILITATOR_URL)

### Approvals (Human-in-the-loop)
- POST /v1/approvals/rules — Add approval rule {"type": "spend_over", "threshold": 100}
- POST /v1/approvals/request — Request human approval for an action
- GET /v1/approvals/{id} — Check approval status
- POST /v1/approvals/{id}/approve — Approve a request
- POST /v1/approvals/{id}/deny — Deny a request
- GET /v1/approvals/pending — List pending approvals

### Proxy
- POST /v1/proxy/upstreams — Register upstream MCP server {"name": "myserver", "url": "https://..."}
- GET /v1/proxy/upstreams — List upstream servers and health
- GET /v1/proxy/tools — List all tools available through proxy
- POST /v1/proxy/call — Call tool through governance proxy {"tool": "scan_domain", "arguments": {...}, "session_id": "ses_xxx"}
- POST /v1/proxy/policies — Add enforcement policy {"type": "block_tool", "tool": "dangerous_tool"}

### Webhooks
- POST /v1/webhooks — Register webhook {"url": "https://hooks.slack.com/xxx", "events": ["anomaly", "approval_requested"]}
- GET /v1/webhooks — List webhooks

### Usage & Metrics
- GET /v1/usage — Usage stats for billing
- GET /v1/metrics — Full platform metrics

## MCP Server

Haldir ships as an MCP stdio server — installable in Claude Desktop, Cursor, Windsurf, and any MCP-compatible client with one config entry:

    pip install haldir
    haldir mcp config  # prints a ready-to-paste mcpServers snippet

Claude Desktop config:

    {"mcpServers": {"haldir": {
        "command": "haldir",
        "args":    ["mcp", "serve"],
        "env":     {"HALDIR_API_KEY": "hld_your_key"}
    }}}

Exposes 19 tools — Haldir's governance primitives: haldir_create_session, haldir_check_permission, haldir_log_audit_action, haldir_get_tree_head, haldir_get_inclusion_proof, haldir_get_consistency_proof, haldir_store_secret, haldir_compliance_score, haldir_build_evidence_pack, and more.

HTTP/SSE endpoint: POST https://haldir.xyz/mcp (JSON-RPC 2.0) — same tool names, 10 of the 19 implemented server-side.
Smithery: https://smithery.ai/server/haldir/haldir
PyPI: https://pypi.org/project/haldir/

## Quick Start

1. Create an API key: POST /v1/keys
2. Create a session: POST /v1/sessions {"agent_id": "my-agent", "scopes": ["read", "spend:50"]}
3. Use the session_id for all subsequent calls
4. Every action is automatically logged to the audit trail
5. Revoke the session when done: DELETE /v1/sessions/{id}

## Links

- Website: https://haldir.xyz
- Live tamper demo: https://haldir.xyz/demo/tamper (click Tamper, watch the Merkle tree reject it)
- Evidence pack: https://haldir.xyz/compliance?demo=1
- API Docs: https://haldir.xyz/docs
- OpenAPI Spec: https://haldir.xyz/openapi.json
- AGENTS.md: https://haldir.xyz/AGENTS.md (conventions for AI coding agents touching this repo)
- THREAT_MODEL.md: https://haldir.xyz/THREAT_MODEL.md (STRIDE analysis + named adversaries + honest residual-risk declarations)
- MCP manifest: https://haldir.xyz/.well-known/mcp/mcp.json
- MCP server card: https://haldir.xyz/.well-known/mcp/server-card.json
- JWKS (STH public keys): https://haldir.xyz/.well-known/jwks.json
- ai-plugin.json: https://haldir.xyz/.well-known/ai-plugin.json
- GitHub: https://github.com/ExposureGuard/haldir
- Smithery: https://smithery.ai/server/haldir/haldir
- PyPI: https://pypi.org/project/haldir/
