Metadata-Version: 2.5
Name: agentbyte
Version: 0.39.0
Summary: A toolkit for designing multiagent systems
Author-email: MrDataPsycho <mr.data.psycho@gmail.com>
License-Expression: LicenseRef-Proprietary
License-File: LICENSE
Requires-Python: >=3.11
Requires-Dist: httpx>=0.28.1
Requires-Dist: pydantic-settings>=2.13.0
Requires-Dist: pydantic>=2.12.5
Requires-Dist: pyyaml>=6.0.3
Provides-Extra: all
Requires-Dist: aioboto3>=15.5.0; extra == 'all'
Requires-Dist: aiosqlite>=0.22.1; extra == 'all'
Requires-Dist: arxiv>=2.1; extra == 'all'
Requires-Dist: asyncpg>=0.31.0; extra == 'all'
Requires-Dist: azure-identity>=1.25.1; extra == 'all'
Requires-Dist: beautifulsoup4>=4.12; extra == 'all'
Requires-Dist: fastapi>=0.135.2; extra == 'all'
Requires-Dist: gepa>=0.1; extra == 'all'
Requires-Dist: graphviz>=0.21; extra == 'all'
Requires-Dist: html2text>=2024.2; extra == 'all'
Requires-Dist: openai>=1.107.1; extra == 'all'
Requires-Dist: opentelemetry-api>=1.39.1; extra == 'all'
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.39.1; extra == 'all'
Requires-Dist: opentelemetry-sdk>=1.39.1; extra == 'all'
Requires-Dist: sqlalchemy[asyncio]>=2.0.43; extra == 'all'
Requires-Dist: sqlmodel>=0.0.38; extra == 'all'
Requires-Dist: tiktoken>=0.9.0; extra == 'all'
Requires-Dist: uvicorn[standard]>=0.44.0; extra == 'all'
Requires-Dist: youtube-transcript-api>=0.6; extra == 'all'
Provides-Extra: aws
Requires-Dist: aioboto3>=15.5.0; extra == 'aws'
Provides-Extra: azureopenai
Requires-Dist: azure-identity>=1.25.1; extra == 'azureopenai'
Requires-Dist: openai>=1.107.1; extra == 'azureopenai'
Provides-Extra: openai
Requires-Dist: openai>=1.107.1; extra == 'openai'
Provides-Extra: optim
Requires-Dist: gepa>=0.1; extra == 'optim'
Provides-Extra: otel
Requires-Dist: opentelemetry-api>=1.39.1; extra == 'otel'
Requires-Dist: opentelemetry-exporter-otlp-proto-http>=1.39.1; extra == 'otel'
Requires-Dist: opentelemetry-sdk>=1.39.1; extra == 'otel'
Provides-Extra: postgres
Requires-Dist: asyncpg>=0.31.0; extra == 'postgres'
Requires-Dist: sqlalchemy[asyncio]>=2.0.43; extra == 'postgres'
Requires-Dist: sqlmodel>=0.0.38; extra == 'postgres'
Provides-Extra: research
Requires-Dist: arxiv>=2.1; extra == 'research'
Requires-Dist: beautifulsoup4>=4.12; extra == 'research'
Requires-Dist: html2text>=2024.2; extra == 'research'
Requires-Dist: youtube-transcript-api>=0.6; extra == 'research'
Provides-Extra: sql
Requires-Dist: aiosqlite>=0.22.1; extra == 'sql'
Requires-Dist: asyncpg>=0.31.0; extra == 'sql'
Requires-Dist: sqlalchemy[asyncio]>=2.0.43; extra == 'sql'
Requires-Dist: sqlmodel>=0.0.38; extra == 'sql'
Provides-Extra: sqlite
Requires-Dist: aiosqlite>=0.22.1; extra == 'sqlite'
Provides-Extra: tiktoken
Requires-Dist: tiktoken>=0.9.0; extra == 'tiktoken'
Provides-Extra: viz
Requires-Dist: graphviz>=0.21; extra == 'viz'
Provides-Extra: webui
Requires-Dist: fastapi>=0.135.2; extra == 'webui'
Requires-Dist: uvicorn[standard]>=0.44.0; extra == 'webui'
Description-Content-Type: text/markdown

# Agentbyte

Agentbyte is an observability-first agentic AI framework for building and studying multiagent systems with a learning-first, implementation-oriented workflow.

Current release: **0.39.0**

## Building an Agent

Every example below builds on the same domain: a customer **support agent**. Start with a plain agent and a model client — no tools, no middleware:

```bash
pip install "agentbyte[openai]"   # the model client lives in an extra; see Install below
```

```python
from agentbyte import Agent
from agentbyte.llm import OpenAIChatCompletionClient

model_client = OpenAIChatCompletionClient.from_api_key(model="gpt-4.1-mini")

support_agent = Agent(
    name="support_agent",
    description="Answers customer support questions about orders and shipping.",
    instructions="You are a helpful support agent. Be concise and accurate.",
    model_client=model_client,
)

response = await support_agent.run("Where is my order #1234?")

print(response.final_message.content)
print(response.usage)          # tokens, cost, cache hits
print(response.finish_reason)  # "stop" | "max_iterations" | ...
```

`run()` executes to completion and returns one `AgentResponse`. For live progress — token-by-token streaming, tool calls as they happen — use `run_stream()` instead, which yields events and finishes with that same `AgentResponse`:

```python
async for item in support_agent.run_stream("Where is my order #1234?", verbose=True):
    print(item)
```

## Adding Tools

A support agent is only useful once it can look things up. Turn any function into a tool with `@tool`; gate risky ones with `approval_mode`:

```python
from agentbyte import Agent
from agentbyte.tools import ApprovalMode, tool

@tool
def get_order_status(order_id: str) -> str:
    """Look up the status of an order."""
    return f"Order {order_id} is in transit."

@tool(approval_mode=ApprovalMode.ALWAYS)
def issue_refund(order_id: str, amount: float) -> str:
    """Issue a refund for an order (requires human approval)."""
    return f"Refunded ${amount} for order {order_id}"

support_agent = Agent(
    name="support_agent",
    description="Answers order/shipping questions and can issue refunds.",
    instructions="Look up orders before answering. Refunds always need approval.",
    model_client=model_client,
    tools=[get_order_status, issue_refund],
)
```

Agentbyte also ships ready-made tools you can drop in without writing any code — pass them straight into `tools=[...]`:

- **Core:** `ThinkTool`, `TaskStatusTool`, `CalculatorTool`, `DateTimeTool`, `JSONParserTool`, `RegexTool` (all at once via `create_core_tools()`)
- **Coding:** `ReadFileTool`, `WriteFileTool`, `ListDirectoryTool`, `GrepSearchTool`, `BashExecuteTool`, `PythonREPLTool`
- **Memory:** `MemoryTool` — lets the agent read/write a memory backend (`ListMemory` or `FileMemory`) as a tool call, on top of the automatic context injection every agent already gets

## Agentic Workflows

When support handling is a fixed multi-step pipeline rather than a single agent call — triage, then route — model it as a `Workflow` instead:

```python
import asyncio
from pydantic import BaseModel
from agentbyte.workflow import FunctionStep, StepMetadata, Workflow, WorkflowConfig, WorkflowRunner

class TicketInput(BaseModel):
    text: str

class TriagedTicket(BaseModel):
    text: str
    priority: str

async def triage(input_data: TicketInput, context) -> TriagedTicket:
    priority = "high" if "urgent" in input_data.text.lower() else "normal"
    return TriagedTicket(text=input_data.text, priority=priority)

async def route(input_data: TriagedTicket, context) -> TriagedTicket:
    return input_data  # e.g. assign to a queue here

workflow = Workflow(WorkflowConfig(name="support_ticket_pipeline"))
workflow.chain(
    FunctionStep("triage", StepMetadata(name="triage"), TicketInput, TriagedTicket, triage),
    FunctionStep("route", StepMetadata(name="route"), TriagedTicket, TriagedTicket, route),
)

execution = asyncio.run(
    WorkflowRunner().run(workflow, {"text": "urgent: order not received"})
)
print(execution.state["route_output"])
```

Steps can also wrap an agent (`AgentStep`), call HTTP endpoints (`HttpStep`), transform data (`TransformStep`), or nest another workflow (`SubWorkflowStep`) — with conditional routing, parallel branches, checkpoint/resume, and human-in-the-loop suspend/resume.

### Choosing Agents or Graph-Based Orchestration

Graph-based orchestration makes routing deterministic: after one defined step completes, the next defined step runs. It does not make an LLM's interpretation, tool selection, or output deterministic. Use a graph or `Workflow` when the business process is known, stable, and requires explicit control over every transition.

Start with an `Agent` when the work requires the system to decide how to solve an open-ended request: researching, interpreting documents, selecting tools, synthesizing results, or handling long-tail exceptions. A single agent can complete many such multi-step tasks without encoding every anticipated reasoning path as graph infrastructure.

Keep deterministic controls at operational boundaries. Validate structured output, apply policy checks and approvals, use idempotent writes, and emit audit records before an agent can produce external side effects. This separates deterministic process control from the model's inherently probabilistic reasoning, while allowing a workflow to be introduced where a stable business process truly needs one.

## Orchestrator Patterns

Two different ways to combine multiple agents — pick based on who's in control:

- **`AgentAsTool`** — one agent decides *if and when* to delegate. The support agent stays the single decision-maker and calls the billing agent like any other tool.
- **An orchestrator** (e.g. `RoundRobinOrchestrator`) — a separate controller drives the conversation between agents in turns, until a termination condition fires. Neither agent decides when the other speaks.

**Agent as a tool** — the support agent delegates billing questions:

```python
from agentbyte import Agent
from agentbyte.agents import AgentAsTool

billing_agent = Agent(
    name="billing_agent",
    description="Handles billing and refund questions.",
    instructions="Answer billing questions and process refund requests.",
    model_client=model_client,
)

support_agent = Agent(
    name="support_agent",
    description="Front-line support agent that can delegate billing issues.",
    instructions="Handle general support; delegate billing questions to the billing tool.",
    model_client=model_client,
    tools=[AgentAsTool(agent=billing_agent)],
)
```

**Orchestrated turns** — support and escalation agents collaborate until the ticket is resolved:

```python
from agentbyte import (
    MaxMessageTermination,
    RoundRobinOrchestrator,
    TextMentionTermination,
    UserMessage,
)

orchestrator = RoundRobinOrchestrator(
    agents=[support_agent, escalation_agent],
    termination=TextMentionTermination("RESOLVED") | MaxMessageTermination(6),
)

task = UserMessage(content="Customer says their package never arrived.", source="user")
async for item in orchestrator.run_stream(task, verbose=True):
    print(item)
```

Other orchestrators follow the same `run()`/`run_stream()` shape: `AIOrchestrator` (a model picks the next speaker), `HandoffOrchestrator` (agents explicitly hand off control), `PlanBasedOrchestrator` (a plan is drafted, then executed step by step).

## Middleware

Built in: `LoggingMiddleware`, `PIIRedactionMiddleware`, `GuardrailMiddleware`, `MetricsMiddleware`, `RateLimitMiddleware`, `ApprovalMiddleware`, `ContextCompactionMiddleware`, `RetryMiddleware`, `OTelMiddleware`. Attach any combination to the same support agent via `middlewares=[...]`:

```python
from agentbyte import Agent
from agentbyte.middleware import ApprovalMiddleware, LoggingMiddleware, RateLimitMiddleware

support_agent = Agent(
    name="support_agent",
    description="Answers order/shipping questions and can issue refunds.",
    instructions="Look up orders before answering. Refunds always need approval.",
    model_client=model_client,
    tools=[get_order_status, issue_refund],
    middlewares=[
        LoggingMiddleware(),
        RateLimitMiddleware(max_requests=10, window_seconds=60),
        ApprovalMiddleware(tool_names=["issue_refund"]),
    ],
)
```

## Observability-First Telemetry

Agentbyte exposes two complementary telemetry layers via `OTelMiddleware`:

- **Per-call spans** (`chat <model>`, `tool <name>`, `embedding <model>`) for model/tool/embedding-level diagnostics.
- **Task-level root span** (`agent <name>`) wrapping every per-call span in one run, carrying the final aggregated usage and outcome.

Enable telemetry:

```bash
export AGENTBYTE_ENABLE_OTEL=true
```

Per-call span attributes emitted by `OTelMiddleware`:

- `gen_ai.system`, `gen_ai.operation.name`, `gen_ai.agent.name`, `gen_ai.session.id`
- `gen_ai.request.model`, `gen_ai.response.model`, `gen_ai.response.finish_reason`
- `gen_ai.usage.input_tokens`, `gen_ai.usage.output_tokens`, `gen_ai.usage.total_tokens`, `gen_ai.usage.cost_estimate_usd`
- `gen_ai.tool.name`, `gen_ai.tool.success`
- `gen_ai.embedding.input_count`, `gen_ai.embedding.output_count`
- `gen_ai.input.messages`, `gen_ai.output.messages`, `gen_ai.tool.parameters`, `gen_ai.tool.result` (opt-in content capture — only when `AGENTBYTE_OTEL_CAPTURE_CONTENT=true`, since these can carry PII)

Reading a trace: `chat gpt-4.1-mini` spans show **per-call** usage/cost/finish reason; the wrapping `agent <name>` span shows the **final accumulated** usage and task outcome for the whole `run()`/`run_stream()` call.

## Installation

Python requirement: **3.11+**

The base package (`pip install agentbyte`) has no model provider or storage driver. Add the
extras for the features you use. Using a feature without its extra raises
`MissingExtraError` (an `ImportError`) that names the extra to install.

```bash
pip install "agentbyte[openai,sqlite]"
# or
uv add "agentbyte[openai,sqlite]"
```

| Extra | Installs | Needed for |
| --- | --- | --- |
| `openai` | `openai` | `OpenAIChatCompletionClient`, `OpenAIEmbeddingClient` |
| `azureopenai` | `openai`, `azure-identity` | Azure OpenAI clients and Entra ID credentials |
| `aws` | `aioboto3` | S3 dataset sources, writers and publishers |
| `otel` | OpenTelemetry API, SDK, OTLP exporter | OpenTelemetry tracing and metrics |
| `webui` | `fastapi`, `uvicorn` | `agentbyte webui` |
| `sqlite` | `aiosqlite` | SQLite datasets, workflow checkpoints, evaluation history |
| `postgres` | `sqlalchemy[asyncio]`, `sqlmodel`, `asyncpg` | `SqlUsageBackend`, `SqlEvalRunStore` |
| `sql` | `sqlite` + `postgres` | Both SQL backends |
| `tiktoken` | `tiktoken` | `TiktokenCounter` for exact compaction token counts |
| `optim` | `gepa` | `optimize_with_gepa()` |
| `viz` | `graphviz` | `WorkflowVisualizer.to_dot()` |
| `research` | `beautifulsoup4`, `html2text`, `arxiv`, `youtube-transcript-api` | `agentbyte.tools.research_tools` |
| `all` | every extra above | Everything |

Contributors working in this repository install every extra plus the local tooling groups
(`test`, `dev`, `experiment`):

```bash
uv sync --all-extras --all-groups
```

## Run The WebUI

The WebUI has two sections: the **Playground** (chat with agents, orchestrators and workflows) and
**Evaluations** (run registered evaluation suites and inspect the results). Install the extra first:

```bash
uv sync --extra webui            # add --extra sqlite for Evaluations
```

### Option 1: Sample apps

No API key needed: three demo agents, a team and a workflow on a fake model client.

```bash
uv run python examples/webui/in_memory.py                          # http://127.0.0.1:8070
uv run python examples/webui/in_memory.py --eval-history eval.sqlite   # plus Evaluations
```

With your OpenAI or Azure OpenAI credentials, the preset agents, teams and workflow:

```bash
uv sync --extra webui --extra openai        # or --extra azureopenai
uv run python examples/webui/presets_webui.py                      # http://127.0.0.1:8080
uv run python examples/webui/presets_webui.py --eval-history eval.sqlite
```

Both launchers accept `--port`, `--no-open`, `--eval-history PATH` and `--eval-catalog CATALOG`
(default: the example suites in `examples/eval/workbench_catalog.py`).

### Option 2: Evaluations for your suites

`agentbyte eval ui` opens the Evaluations pages on their own (no Playground). Point it at the
Python file that defines your `EvalSuiteRegistry`:

```bash
uv sync --extra webui --extra sqlite
uv run agentbyte eval ui --catalog examples/eval/workbench_catalog.py      # http://127.0.0.1:8090
```

- `--catalog` takes a `.py` file or a module path (`myapp.evaluations`); add `:name` only when the
  file defines more than one registry. Without `--catalog` it serves the built-in smoke suite.
- Runs are stored in `./eval.sqlite` unless you pass `--history PATH`. Only one process can write
  a history file: stop the UI before `agentbyte eval run` on the same file, or use another file.
- The same catalog works in the terminal: `uv run agentbyte eval run capitals --catalog examples/eval/workbench_catalog.py`.
- To serve agents and Evaluations together: `agentbyte webui --dir my_entities --eval-catalog myapp/evaluations.py`.

Start with the end-to-end guide, `docs/guides/evaluation-end-to-end.md` (dataset → suite → history → UI), and the
notebook `notebooks/usecases/eval/06-eval-end-to-end.ipynb` (real OpenAI/Azure model). Details: `docs/design-guide/05-eval_topics/`.

### Option 3: Scan a project directory

`agentbyte webui --dir PATH` loads entities from Python files. Discovery is convention-based: each
module (or package) must define a top-level variable named `agent`, `orchestrator` or `workflow`
holding an Agentbyte object. Folders without those names show "No agents, orchestrators, or
workflows found"; this is the case for `examples/`, which uses the launchers above instead.

```python
# my_entities/support.py
agent = Agent(name="support", ...)
```

```bash
uv run agentbyte webui --dir my_entities
uv run agentbyte webui --dir my_entities --eval-catalog myapp/evaluations.py   # both sections
```

### Option 4: From Python

```python
from agentbyte.webui import serve

serve(entities=[agent], port=8080)
serve(entities=[agent], eval_registry=registry, eval_history="eval.sqlite")   # with Evaluations
```

### Quick Troubleshooting

- Port in use: add `--port 8090`. The UI works on any port.
- No browser wanted: add `--no-open`.
- "No agents, orchestrators, or workflows found": the `--dir` folder has no module with a top-level
  `agent`, `orchestrator` or `workflow`; use a launcher from Option 1 to see the samples.
- No Evaluations switch in the header: start `agentbyte webui` with `--eval-catalog`, or use `agentbyte eval ui`.
- After upgrading, hard-reload the page (Cmd/Ctrl-Shift-R) to load the new UI.

## Development

```bash
uv run ruff check src tests
uv run pytest tests -v
```
