Metadata-Version: 2.4
Name: suspense
Version: 0.4.0
Summary: The fuse box for AI agents: freeze, inspect, resume or stop a runaway agent run without losing its state.
License: Apache-2.0
Keywords: llm,agents,proxy,circuit-breaker,observability
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Operating System :: OS Independent
Classifier: Topic :: System :: Monitoring
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pyyaml>=6
Provides-Extra: fast
Requires-Dist: httpx>=0.27; extra == "fast"
Dynamic: license-file

# Suspense — the fuse box for AI agents

[![ci](https://github.com/zlik/suspense/actions/workflows/ci.yml/badge.svg)](https://github.com/zlik/suspense/actions/workflows/ci.yml)
[![PyPI](https://img.shields.io/pypi/v/suspense)](https://pypi.org/project/suspense/)
[![license](https://img.shields.io/badge/license-Apache--2.0-blue)](LICENSE)

Agents run unattended. When one goes wrong (loops forever, burns money, does something
you didn't expect) your only options today are *watch it* or *kill it*. Killing it loses
everything it was doing.

Suspense gives you a third option: **freeze it, look at it, then resume or stop it.**

It's a tiny HTTP proxy that sits between your agent and its model provider. Because every
model call carries the agent's full working state (the message array), Suspense can meter
every step, trip a breaker when a run crosses a limit, and hold the agent simply by not
answering yet. The agent blocks. Nothing is lost. An operator decides what happens next.

No framework integration. No code changes. Point `base_url` at Suspense and you're done.

Three things that go wrong with unattended agents, and what happens instead: one that
should stop itself, one that needs a human before an irreversible action, and one whose
work must not be lost. This is `examples/demo.sh`, offline, forty seconds, real agent code
with a scripted model:

![demo: the overnight bill, the refund, nothing is lost](docs/demo.svg)

## Quick start

```bash
pipx install suspense
suspense init                      # writes suspense.yaml with a token, prints the one-line change for your SDK
suspense serve                     # :4141; /v1/messages -> Anthropic, everything else -> OpenAI
```

In your agent, change one line:

```python
client = OpenAI(base_url="http://localhost:4141/v1")        # OpenAI SDK (Chat Completions and Responses)
client = Anthropic(base_url="http://localhost:4141")        # Anthropic SDK
```

The agents in the demo are real code you can copy: a [support agent](examples/support_agent.py) that
refunds and emails, a [research agent](examples/research_agent.py) that gets stuck, and a
[migration agent](examples/migration_agent.py) that must not lose its place. The same one-line
change works for [LangChain](examples/langchain_agent.py), [LangGraph](examples/langgraph_agent.py),
the [OpenAI Agents SDK](examples/agents_sdk_agent.py) and [Node](examples/node_agent.mjs).

Then operate, from the terminal or the fleet view at `http://localhost:4141/suspense/`:

```bash
suspense ls                                  # every run, status, steps, cost, why it's held
suspense show <run> 14                       # the exact request at step 14: the agent's full state
suspense resume <run> --steps 20 --cost 1.50 # continue with a budget you choose
suspense stop <run>                          # the agent's next call gets a 409 and exits cleanly
suspense hold --tag prod                     # freeze everything under a tag during an incident
```

## The three scenarios

**The overnight bill.** A research agent is asked for a digest at 6pm and left running. It
reads two pages, then searches for "the latest" again, and again; every step re-reads the
whole conversation, so each one costs more than the last. Nobody wants to be paged for
this, so nobody is: the third identical search is a loop, Suspense ends the run itself at
step five, and Slack gets one line for whoever reads it in the morning, with the record of
what the agent was doing. Cost of the incident: a few cents. The cost cap is the backstop
for runaways that don't repeat themselves, and the budget per team is the backstop for
everything else. The bill is attributed to the agent and the team, not discovered on an
invoice.

**The refund.** A support agent handles "order #48213 arrived damaged". Company policy says
refunds over $100 need a human; the agent's code doesn't know that. The model decides to
refund $349 and Suspense holds that answer before the agent sees it. Deny, and the agent is
told the refund was not performed and is under review, so it tells the customer exactly
that and closes out the ticket; if nobody decides within fifteen minutes, the same thing
happens on its own. The same agent refunds $45 on the next ticket without anyone noticing.
And the Deny is now a test: replay it on a cheaper model or an edited prompt and find out
whether the policy would have held without the proxy.

**Nothing is lost.** A migration agent is working through a batch. An incident: freeze
everything tagged `prod`. The agent is alive and parked mid-batch with its state intact.
Resume, and it continues from the next record. The alternative was kill, and start over.

## What it does

- **Breakers per run**: cost, steps, minutes, tokens per minute. Evaluated *before* each call, so a tripped run never spends another cent. Per-tag limits and budgets across runs cap the bill, not just the worst run.
- **Hold instead of kill.** A held agent is a waiting agent: Suspense keeps its SDK waiting past its own timeout, on every shape (OpenAI Chat Completions, OpenAI Responses, Anthropic Messages), streaming or not.
- **Tool-call gate.** When the model asks for `delete_*`, or `send_email` to an outside domain, the run is held *before the agent sees the answer*. Approve or Deny from the fleet view or Slack; if nobody decides in fifteen minutes, policy decides. A denial reaches the agent as a normal answer, "not performed, under review, finish up," so the flow completes instead of crashing.
- **Nobody has to watch.** Trips stop the run and say why, rate limits wait themselves out, spent budgets refuse the next call, loops end on the spot. Holding is opt-in per tag, for work worth a person's time.
- **Loop detection.** The same tool with the same arguments N times in a row is held long before a step cap would fire.
- **Every step is a snapshot.** Inspect it, export the run as a forensic bundle, fork it with an edited prompt, or re-run it on a cheaper model and diff the answers.
- **Where the money goes.** `suspense insights` reads the snapshots across runs, models and steps and says what's wasting money: repeated work, the context tax of re-sending the conversation, a cheaper model that finishes as often, runs that never finish. Each finding comes with the eval command that proves the fix. Nothing runs in the agent's path.
- **Every decision is a test.** Approve, Deny and loop holds become eval cases at the exact state they happened; `suspense evals run --model X` replays them and tells you what a model or prompt change broke.

## What it fixes

Agents are the first software that spends money and takes actions on its own, at a pace
set by a model rather than a person. The engineering symptom is "it loops". The business
problems are what the loop does to a budget, to work already done, to the systems the
agent can touch, and to whoever answers for it afterwards. Suspense is aimed at four of them.

| The pain | Who feels it | What changes with Suspense |
|---|---|---|
| **Unbounded spend.** A run has no natural ceiling, and each step re-reads the whole conversation so cost per step grows. The provider bill arrives the next day with no idea which run caused it. | Finance, platform lead | A hard ceiling per run in dollars, steps, minutes and burn rate, enforced *before* the next call. Spend attributed to a run, a tag, a team. |
| **Killing loses the work.** Today the only way to stop a bad run is to kill it. Hours of tool calls and paid tokens are gone, and the re-run often repeats the mistake. | Agent team, operations | Hold instead of kill. The agent keeps its full state; resume continues from that exact call. Nothing is redone or re-bought, and you see the state before deciding. |
| **Irreversible actions with no checkpoint.** An agent that can delete, send, pay or deploy will eventually be told to by its own model. After-the-fact review doesn't prevent it. | Security, compliance | The tool-call gate puts a human between the model's decision and its execution, for exactly the tools you name. Every step is a snapshot, exportable as one file for the incident review. |
| **Nobody can see the fleet.** Past a handful of agents, "what is running, what is it costing per minute, can I stop the one that matters?" has no answer short of grepping logs. | Engineering manager, on-call | One page, every run, live, sorted by dollars per minute, with hold, resume and stop per row. Freeze everything tagged `prod` in one command. Slack alert with the reason. |

The impact is measurable per team with three numbers: runaway runs per month, what each
cost before someone noticed, and the engineer hours spent noticing, killing and re-running.
The breakers remove the first, hold-and-resume most of the second, the fleet view and alerts
shrink the third. The number that decides the security conversation, the cost of one
irreversible action, is the one no team can quote in advance.

**Who it's for.** The platform lead who owns the provider bill and wants a ceiling and
attribution. The security lead asked to sign off on agents touching production, who needs a
checkpoint before actions and evidence after, and who can read the whole thing in one file.
The team shipping the agent, who wants to leave it running overnight and deal with a trip in
the morning without losing the night's work.

**What it is not.** It only sees model calls that go through it, and a hold lands at the
next call, not mid-tool. It is one proxy and one SQLite file, not a control plane. It stores
prompts and responses locally so you can inspect and replay them; that is the feature, and
retention is yours to set.

## Not a gateway

Gateways route, hold API keys, split traffic, filter content, and reject requests that
exceed a budget. They are the right place for all of that, and Suspense sits in front of
one happily: set `upstream` to the gateway and keep everything else. What a gateway
can't do is the reason Suspense exists: it doesn't know which requests belong to one
agent's run, so it can't cap that run; when a limit hits, it can only refuse, which kills
the run instead of parking it; and it never withholds a model's answer so a human can
approve the tool call inside it. Suspense is the intervention layer. It acts on the run,
not the endpoint, and it holds instead of rejecting.

## Verified against the real thing

The one-line promise is tested in CI against the real OpenAI and Anthropic SDKs for Python
and Node, LangChain, LangGraph and the OpenAI Agents SDK, held past their timeouts and
resumed. The server holds 2,000 parked agents on one thread. See
[docs/reference.md](docs/reference.md) for configuration, run identification, the gate,
evals, export and retention, and what it doesn't do.

## License and what stays free

Suspense is Apache-2.0. Everything in this repo is and stays free: the proxy, the breakers,
hold/resume/stop, the tool-call gate, selectors, the fleet view, export, replay, evals. A
single proxy should be something a security team can read in an afternoon and run anywhere:
fourteen small modules, no runtime dependencies.

What will cost money, later and in a separate product: the things one proxy can't do.
A control plane across many proxies and teams, SSO and roles, an audit log of who
resumed what, approvals from Slack, hosted retention and cost reporting over time.
