Metadata-Version: 2.5
Name: timefork
Version: 0.2.0
Summary: git bisect for agent runs. Record once through a proxy, replay for free, change one thing, find the step that broke it.
Author: Arun Shankar
License: Apache-2.0
License-File: LICENSE
License-File: NOTICE
Keywords: agents,debugging,llm,mcp,observability,replay
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Software Development :: Debuggers
Requires-Python: >=3.11
Requires-Dist: click>=8.1
Requires-Dist: httpx>=0.27
Requires-Dist: rich>=13.9
Requires-Dist: starlette>=0.41
Requires-Dist: uvicorn>=0.32
Provides-Extra: aws
Requires-Dist: botocore>=1.34; extra == 'aws'
Provides-Extra: dev
Requires-Dist: mypy>=1.13; extra == 'dev'
Requires-Dist: pytest-asyncio>=0.24; extra == 'dev'
Requires-Dist: pytest>=8.3; extra == 'dev'
Requires-Dist: ruff>=0.7; extra == 'dev'
Provides-Extra: frameworks
Requires-Dist: google-adk>=1.0; extra == 'frameworks'
Requires-Dist: langchain-anthropic>=0.3; extra == 'frameworks'
Requires-Dist: langchain-openai>=0.2; extra == 'frameworks'
Requires-Dist: langgraph>=0.2; extra == 'frameworks'
Requires-Dist: litellm>=1.40; extra == 'frameworks'
Requires-Dist: openai-agents>=0.1; extra == 'frameworks'
Requires-Dist: openai>=1.40; extra == 'frameworks'
Requires-Dist: pydantic-ai>=1.0; extra == 'frameworks'
Requires-Dist: strands-agents>=1.0; extra == 'frameworks'
Description-Content-Type: text/markdown

<h1 align="center">Timefork</h1>

<p align="center">
  <b>git bisect for agent runs.</b><br>
  Record a run once, through a proxy. Replay it for free. Change one thing. Find the step that broke it.
</p>

<p align="center">
  <a href="#install"><img alt="python" src="https://img.shields.io/badge/python-3.11%2B-1E2838?style=flat-square&labelColor=0F1520"></a>
  <a href="LICENSE"><img alt="license" src="https://img.shields.io/badge/license-Apache--2.0-2DD4A7?style=flat-square&labelColor=0F1520"></a>
  <a href="#it-works-with-your-framework-because-it-has-never-heard-of-it"><img alt="frameworks" src="https://img.shields.io/badge/frameworks-all%20of%20them-8AB4FF?style=flat-square&labelColor=0F1520"></a>
  <a href="#status"><img alt="tests" src="https://img.shields.io/badge/tests-real%20HTTP%20over%20TCP-2DD4A7?style=flat-square&labelColor=0F1520"></a>
  <a href="docs/REAL.md"><img alt="real api" src="https://img.shields.io/badge/proven%20on-the%20real%20API-FFA94D?style=flat-square&labelColor=0F1520"></a>
</p>

![A bisected agent run: the step that broke it, the probes that found it, and the answer that would have worked](docs/images/investigation-light.png)

<p align="center"><sub>A finished investigation. Five steps, three agents, one worker answered badly once. The search was told nothing and named the step in four probes.</sub></p>

<details>
<summary><b>Ten seconds of the timeline</b> - the investigation, the wiring, a what-if, and a real Claude Code session with a subagent</summary>

![Four frames: a bisected run, the wiring of a fan-out, a what-if edit forking at the critic, and a real Claude Code subagent session](docs/images/tour.gif)

</details>

---

```console
$ timefork record -- python debug_agent.py
recorded 8 steps · 41,204 in / 3,180 out · run eb1d7c9dec03

$ vim prompts/system.md          # change one line

$ timefork replay last -- python debug_agent.py
8 steps · 4 from tape · 4 live
diverged at step 4
  step 4: messages[4].content[0].text: "propose a fix" -> "propose a minimal fix"
branch 606ec691ab01
```

**Four of those steps cost nothing and took no time.** Only the ones downstream
of your edit actually ran. You did not say where you edited something; Timefork
worked it out.

<br>

## Why

An agent does a long job in many steps. When it fails at step 30, you cannot
tell which earlier step was to blame. To test a fix you re-run the whole thing
from step 1, which costs money, takes minutes, and may behave differently
anyway. So nobody tests anything. Everyone guesses.

Every observability tool on the market will show you the trace.
**None of them let you fork one and re-execute from the middle.**

<br>

## Install

Nothing to install first if you have [uv](https://docs.astral.sh/uv/):

```bash
uvx timefork record -- claude -p "what does this repo do"
uvx timefork serve
uvx timefork replay last -- claude -p "what does this repo do"    # free
```

Measured cold on a laptop with a one-line prompt: the first command was
running 1.2 s after being typed, the Claude Code run recorded in 9 s, the
page was up 1.5 s later, and the replay finished in 2.5 s without touching
the provider. For keeps:

```bash
pip install timefork        # or: uv tool install timefork
cd your-project && timefork init
timefork record -- python my_agent.py
timefork serve
```

No agent code changes. No SDK to import. No account. Try it with no key at
all: `python examples/seed_demo.py && timefork serve` records the bundled
multi-agent pipeline with one worker scripted to fail, bisects it, and opens
the finished investigation above.

<br>

## Three things it does that nothing else does

### 1. Replay for free, and fork exactly where you changed something

```
             0     1     2     3     4     5     6     7
  baseline  ████  ████  ████  ████  ████  ████  ████  ████
  branch    ░░░░  ░░░░  ░░░░  ░░░░  ████  ████  ████  ████
                                    ◆
                                    fork · step 4

  ░░░░  served from tape · free, instant, no provider touched
  ████  executed live · billed
  ◆     the first call whose request changed
```

Every call is filed under the **hash of its request**, not its position. Edit
a prompt and replay: every call before your edit hashes identically and is
served from the tape; the first call that misses *is* the fork. Three workers
running concurrently finish in a different order every run, and it does not
matter, because each is filed under its own question.

### 2. Find the step that broke it

```console
$ timefork bisect last --check "pytest -q"
bisecting eb1d7c9dec03 · 8 steps · binary search

  replay unchanged       fail   reproduce
  live from step 0       pass   resample all
  live from step 4       pass
  live from step 6       fail
  live from step 5       fail

step 4 is the step that broke it
  claude-demo-1  Use decimal.Decimal with ROUND_HALF_UP instead of int().
step 4 is where the run became unsalvageable. Resampling from there
rescued it; resampling from 5 did not.
```

`probe(i)` replays with everything before `i` from the tape and everything
from `i` on answered freshly: *if the model had answered differently from here,
would the run have survived?* Binary search finds the last step where
intervening still worked, in about `log2(n)` partial re-runs. Two verdicts are
more useful than a step number, and it reports them first: a pure replay that
**passes** means the failure is not in the model calls at all; resampling
everything that still **fails** means the agent is reliably wrong, not unlucky.

Most agent failures are "the answer is wrong", which no exit code knows.
`--judge "the answer names the largest row"` has a model read the agent's
output and answer PASS or FAIL - one short call per probe, never through the
proxy - so the search applies to any run at all, test or no test.

On a fault-injection benchmark with the broken step known, it locates
transient faults exactly **100%** of the time in a mean of 5.7 probes. For
intermittent failures, `-r 4` combines draws by disjunction rather than
majority, which takes exact localisation from 0.12 to 0.84 where resampling
rescues the run half the time. All three verdicts have been confirmed on real
Claude Code sessions, and a real refusal on the real API was called
*deterministic* correctly. `bench/faults.py` measures it against ground
truth.

### 3. Change what the model said, and replay from there

![A worker's answer changed from 66 to 999 on the timeline; three steps replayed free, the critic re-ran live](docs/images/scenario-what-if-light.jpg)

Pick a step, edit its answer - or the tool it called - and replay. The edit is
served to the agent in the vendor's own wire shape; every step that did not
depend on it replays free, and the first one that did is where the run forks.
One question asked of the recording, one step paid for.

```bash
timefork replay REF --edit '3=The largest is delta.'
timefork replay REF --call '0=read_ledger {"name": "ledger-b.txt"}'
```

<br>

## It works with your framework because it has never heard of it

Timefork is an HTTP proxy, not an integration. It sits between your agent and
its model provider. There is nothing to import, subclass, wrap or configure.

| | |
|:--|:--|
| Google ADK | ✓ [on Claude and on Gemini](docs/FRAMEWORKS.md) |
| LangChain · LangGraph | ✓ [on Claude and on OpenAI](docs/FRAMEWORKS.md) · [pictured](docs/SCENARIOS.md) |
| OpenAI Agents SDK, handoffs and tools | ✓ [on OpenAI and on Claude](docs/FRAMEWORKS.md) · [pictured](docs/SCENARIOS.md) |
| CrewAI | ✓ [on Claude and on OpenAI](docs/FRAMEWORKS.md) |
| AutoGen | ✓ [on Claude and on OpenAI](docs/FRAMEWORKS.md) |
| Pydantic AI | ✓ [on Claude](docs/FRAMEWORKS.md) |
| Strands Agents (AWS) | ✓ [on Claude](docs/FRAMEWORKS.md) |
| Claude Code | ✓ real sessions, [pictured](docs/images/subagents-light.jpg) |
| Claude Agent SDK, subagent and all | ✓ real session, [pictured](docs/images/agent-sdk-light.jpg) |
| forty lines of `httpx` | ✓ most of the test suite |
| Cursor · Cline | the same wire shape, not run here: they are desktop programs |
| the thing you wrote this morning | ✓ if it speaks HTTP to a provider |

Every mark above is a program in this repository that was run through the
proxy against a real model, recorded, and then replayed with every step from
the tape and the provider never contacted. The numbers are in
[docs/FRAMEWORKS.md](docs/FRAMEWORKS.md), rebuilt by `python docs/real_frameworks.py`.

If it speaks HTTP to Anthropic, OpenAI, Gemini, Vertex or Bedrock - or to any of
the many APIs that copied the OpenAI shape, including Azure, Groq, Together,
vLLM and Ollama - it is recorded the same way. Anthropic, OpenAI and Gemini on
Vertex AI have each met real traffic here; Bedrock and the OpenAI-shaped clones
are handled by the same adapters and have been run only against fakes.

Bedrock is the one that needs something extra. SigV4 covers the host header, so
the signature your agent produced is for the proxy and AWS rejects it; the proxy
strips it and signs again, which means **the proxy needs its own AWS
credentials**. Install with `pip install 'timefork[aws]'`. The re-signing is
tested against a fake; a real Bedrock run has not been made here.

**Wrap a command:**

```bash
timefork record -- adk run ./my_agent
timefork record -- python -m my_service
```

**Or run standalone**, for servers and containers you cannot wrap:

```bash
timefork up --port 7777
eval "$(timefork env --port 7777)"      # in the other shell
```

<br>

## Proven on the real API

Every scenario in the gallery, run against the real Claude API through the
official SDK - streaming, adaptive thinking, prompt caching, tool loops,
delegation, an image, a PDF, a twelve-turn conversation, a real MCP server
behind Claude Code - and read back off the tape. The full table with every run
is [docs/REAL.md](docs/REAL.md); it is produced by `python docs/real.py` and
rebuilt before a release.

What only real traffic showed, each of which the scripted model could never
have produced: a mid-stream safety refusal on a ledger-adding tool call, which
`bisect` correctly called deterministic; a worker that thought for sixteen
tokens and said nothing; a model delegating to two helpers in one turn; and a
proxy that wrapped an upstream 400 in a 200 event stream. All four are handled
and written up in [docs/REAL.md](docs/REAL.md).

<br>

## Who did what, and who handed what to whom

An agent run is rarely one agent. The timeline tells them apart - the main
agent, each subagent, each MCP server - from the `X-Timefork-Actor` header a
framework may send, or from the system prompt otherwise.

A team of agents is a graph, not a list: a planner fans out to workers, workers
fan in to a critic, a supervisor spawns a subagent that spawns another. Timefork
draws that graph from the traffic alone. When something one call *said* turns
up inside a later call's *question*, that is an edge, and the words that carried
it are on the edge.

```
$ timefork graph last
a993b4f2bd89 subagents
3 threads, 4 handoffs

supervisor  #0 #4
  └ researcher  #1 #3
    └ reader  #2

#0 supervisor → #1 researcher   "Find the largest value in the ledger and say wh…"
#1 researcher → #2 reader       "Read the ledger and list every row with its val…"
#2 reader → #3 researcher       "alpha 12 beta 7 gamma 19 delta 3 epsilon 25"
#3 researcher → #4 supervisor   "The largest is epsilon at 25, on the fifth row."
```

The timeline's **Wiring** view draws the same thing: one lane per conversation,
indented under the one that started it, a curve for every handoff, and after a
bisect, the path the blamed answer took in rose. This is a real Claude Code
session that was asked to spawn a subagent; nothing was declared:

![A real Claude Code session with a subagent, drawn as lanes: the main agent, the subagent nested under it, the brief going down and the report coming back](docs/images/subagents-light.jpg)

Measured against runs with a known graph (`bench/lineage.py`): no errors on
verbatim or paraphrased quotes, 0.99 precision on one-word answers, and an
honest floor where identical short answers cannot be told apart.

<br>

## Tool calls are on the tape too

A model call is half of what an agent does. The other half is tool calls, and
agents fail in those at least as often. Put the shim in front of an MCP server
and both halves land on one timeline:

```json
"command": "timefork",
"args": ["mcp", "--", "npx", "-y", "@modelcontextprotocol/server-filesystem", "."]
```

On replay a recorded tool call is answered from the tape and **the real server
is never asked** - a replayed agent does not read the file again, does not hit
the API again, and does not depend on the world having stayed still. With
nothing recording, the shim execs the real server and disappears, so the config
is safe to leave in place permanently.

<br>

## A test suite you already have

```console
$ timefork ci
checking 2 recordings · block on divergence

  pass   checkout-flow            3 steps from tape
  drift  refund-flow              step 1 (refund agent): "issue the refund" -> "issue a partial refund"

1 unchanged · 1 drifted
4 steps replayed · nothing billed
```

Snapshot testing for agents. Every recording is a snapshot of how the agent
behaved on real input, and `ci` asks whether the prompts, skills and tool
descriptions in your working tree still produce it. **A passing check is
free** - with divergence blocked, a matching replay never reaches a provider.
Red names the step, the agent whose lane it landed in, and the exact
difference, and exits non-zero.

A recording becomes a test with one line: `timefork expect checkout-flow
"order 4471 confirmed"`. Then a drift whose final answer still says that is a
*drift* - the prompts moved, the outcome did not - and one whose answer no
longer says it is a *fail*, which is the thing a regression suite is for. Run
`ci --on-diverge live` to let a drifted run reach its answer and be scored, and
`ci --judge "the order was confirmed"` to have a model score it instead of a
substring.

<br>

## From Python, with callbacks

The command line drives a subprocess. A program can drive its own agent in
the same process, sync or async, and be told what happens as it happens.

```python
from timefork import Timefork

deck = Timefork.open()                              # finds .timefork, or makes one

with deck.record(label="nightly") as take:          # a proxy on a free port
    my_agent(env=take.env)                          # the env is the whole integration
print(take.stats.steps, "steps recorded")

with deck.replay("nightly", on_divergence=print) as take:
    my_agent(env=take.env)
print(take.stats.from_tape, "free")

async with deck.replay("nightly", edits={3: "The largest is delta."}) as take:
    async for kind, event in take.events():         # ("step", Step) or ("divergence", Divergence)
        ...

result = deck.bisect("nightly", play=lambda env: my_agent(env=env) == 0)
```

Every take is a run on the tape, exactly as the command line would have made
it: the timeline shows it, `share` exports it, `bisect` can search it.


An async client that cannot be pointed at a base URL by its environment can
be handed the recorder directly: `httpx.AsyncClient(transport=take.transport())`
reaches it in-process, with no socket. Sync clients take `base_url=take.base_url`.

<br>

## Send someone the recording

```bash
timefork share last -o bug-report.html
```

One file. The viewer with the recording folded into it - no install, no account,
no server, and it opens offline from a downloads folder. A nine-step run with
its parent is about 40 KiB, because payloads are stored once and referenced by
address. Credentials are scrubbed on the way out by default, and the file says
on its face whether it was redacted.

<br>

## The timeline

```bash
timefork serve
```

Two branches as two tracks, aligned step for step. Click a step to see its
request and response; arrow keys scrub; a recording in progress streams in
live. Light and dark. No build step, no Node, no network - the whole front end
ships as one file inside the package.

Two buttons on every run do what the page would otherwise tell you to type:
**Replay** runs the recorded command again under replay and lands you on the
new branch, and **Find the step that broke it** runs the bisect and puts the
finding on the run. Under every answer: **Change what it said, replay from
here.** Each is exactly the command line, run as a process in the directory the
run came from, for the local machine only.

<br>

## Point it at your agent once

Real agents put things in their prompts that change every time they start: a
session id, a timestamp, a working directory, a freshly minted subagent id. None
are changes to the agent, all change the request, so replays diverge for no
reason until somebody works out what to disregard.

Don't work it out. Run the work twice and let it diff itself:

```console
$ timefork calibrate --apply -- python my_agent.py
pass 1 of 2 · recording a39e3d2f5b86
pass 2 of 2 · recording 1cfd5e2449d0

varied across 2 steps:
  metadata.user_id   every step
  system[2].text     1 steps

added 2 · 2 paths ignored for this tape
```

Calibrate with the workload you actually intend to replay; a trivial prompt only
exercises the fields a trivial prompt touches. Things that vary *inside* a
field - an identifier in a sentence - are normalised rather than dropped:

```bash
timefork ignore --ids                 # uuids and hex identifiers
timefork ignore -p 'run-[0-9]+'       # your own varying values
```

Changing the list re-files every recording already on the tape under the new
rules, so ignore the path a divergence named, replay again, and the steps that
only differed there are served from the tape.

**How far this gets you**, measured on real Claude Code sessions: short
sessions replay whole; long ones replay until they meet something the first run
changed in the world. An eleven-step code review served eight steps from the
tape and then stopped exactly where its own `Write` had created a file the
replay now found already existing:

```console
$ timefork replay last --on-diverge block -- claude -p "..."
8 steps · 8 from tape · 0 live
diverged at step 8
  content: "File created successfully at: /private/tmp/…"
        -> "<tool_use_error>File has not been read yet. Read it first…"
```

You can un-watch a video; you cannot un-send an email. That is the limitation,
and it is also the argument for the design: Timefork did not hand back a subtly
wrong replay. A miss is always loud. It will never quietly serve a nearby
cached response, because one silently wrong result would make every other
result untrustworthy. `--on-diverge block` refuses to touch a provider at all,
which is the right default in CI.

<br>

## Commands

| | |
|:--|:--|
| `timefork init` | create a tape here |
| `timefork doctor` | check this machine before the first recording |
| `timefork record -- CMD` | run CMD with recording on |
| `timefork replay REF -- CMD` | replay a run as a new branch; `--edit`, `--call`, `--from`, `--on-diverge block`, `--pace` |
| `timefork bisect REF --check CMD` | find the step that caused a failure; `--judge "…"` lets a model decide; `--ancestors` probes only the steps that reached the last one |
| `timefork serve` | open the timeline |
| `timefork graph REF` | draw who handed what to whom |
| `timefork mcp -- SERVER` | record an MCP server's tool calls |
| `timefork share REF` | export a run as one shareable file |
| `timefork ci` | replay every recording and report what changed; `--judge "…"` lets a model score the answers |
| `timefork expect REF TEXT` | what a recording's final answer must say, for `ci` to score |
| `timefork suspects` | across many recordings, the step whose words predict failure |
| `timefork sweep REF --vary F --variants D` | try many versions of a prompt |
| `timefork calibrate -- CMD` | run twice, find what varies on its own |
| `timefork ignore PATH` | stop a varying field counting as a change |
| `timefork rehash` | re-file the tape after editing its ignore list by hand |
| `timefork redact -p REGEX --path PATH` | never store these values; scrubbed before hashing, so replay still matches; `--hash` keeps a stable token per value |
| `timefork log` · `show REF` · `grep TEXT` · `stat` | list runs, list steps, search, size |
| `timefork up` · `timefork env` | standalone proxy |

`REF` is a run id, a unique prefix, a unique label, or `last`.

<br>

## Design

Four layers. The dependency arrows only ever point downward.

```
  cli · ui · api  user-facing surfaces
  proxy           records and replays HTTP traffic
  providers       per-vendor request and response shapes (pure functions)
  tape            the storage format (knows nothing about HTTP)
```

Every boundary is a `Protocol`. `BlobStore` becomes S3, GCS or Azure Blob.
`Index` becomes Postgres or Spanner. A provider is one file and one registry
entry. Nothing above a layer learns that anything below it changed.
[ARCHITECTURE.md](ARCHITECTURE.md) explains it in plain English and ten
diagrams.

<br>

## Status

`v0.2` - alpha, and honest about it. What is in it: [CHANGELOG.md](CHANGELOG.md).
What is left, and where it is going: [TODO.md](TODO.md).

| | |
|:--|:--|
| Anthropic · OpenAI-shaped · Google | ✓ recorded and replayed |
| AWS Bedrock, re-signed at the proxy | ✓ `pip install 'timefork[aws]'` |
| streaming and unary | ✓ captured verbatim |
| divergence detection and diffing | ✓ |
| `bisect` causal root-cause | ✓ |
| edit-and-replay: words or a tool call | ✓ |
| MCP tool calls, recorded and replayed | ✓ |
| every agent and MCP server told apart on the timeline | ✓ |
| the graph of who handed what to whom, read off the traffic | ✓ |
| shareable single-file export, redacted | ✓ |
| `ci` regression replay over a whole corpus | ✓ |
| Python API, sync and async, with callbacks | ✓ |
| twelve framework runs on Claude, OpenAI and Gemini, none of them knowing | ✓ [docs/FRAMEWORKS.md](docs/FRAMEWORKS.md) |
| cost in dollars rather than tokens | ✗ |

Every test runs a real upstream, a real proxy and real HTTP over TCP. The
assertion that matters most is that a replay leaves the provider's call count
untouched.

<br>

## Read more

- [docs/SCENARIOS.md](docs/SCENARIOS.md) - every scenario the tool is built for, as a real tape with a picture, in both themes, rebuilt by a script.
- [docs/REAL.md](docs/REAL.md) - the same scenarios on the real API, with numbers.
- [docs/FRAMEWORKS.md](docs/FRAMEWORKS.md) - seven agent frameworks recorded and replayed on the real API, none of them knowing.
- [docs/LAUNCH.md](docs/LAUNCH.md) and [docs/VIDEO.md](docs/VIDEO.md) - the launch post and the sixty-second video's script, both built from the numbers above.
- [ARCHITECTURE.md](ARCHITECTURE.md) - how it works, in plain English and diagrams.
- [SECURITY.md](SECURITY.md) - what the tape holds, what `share` scrubs, what the page can run, and `timefork redact` for values a tape must never hold.

## License

Apache 2.0. Contributions are DCO-signed - see [CONTRIBUTING.md](CONTRIBUTING.md).
Timefork is a trademark of its maintainer; see [NOTICE](NOTICE).
