Metadata-Version: 2.5
Name: agentcore-dashboard-metrics
Version: 0.6.0
Summary: Plug-and-play turn-metrics logging for LangGraph/Strands/CrewAI agents on Bedrock AgentCore, shaped for CloudWatch Logs Insights + Grafana dashboards.
Author-email: nakulan <nakult721@gmail.com>
License-Expression: MIT
License-File: LICENSE
Keywords: agentcore,bedrock,cloudwatch,crewai,grafana,langchain,langgraph,logging,observability,strands-agents
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: System :: Logging
Classifier: Topic :: System :: Monitoring
Classifier: Typing :: Typed
Requires-Python: >=3.9
Provides-Extra: dev
Requires-Dist: build; extra == 'dev'
Requires-Dist: pytest; extra == 'dev'
Requires-Dist: twine; extra == 'dev'
Description-Content-Type: text/markdown

# agentcore-dashboard-metrics

Plug-and-play turn-metrics logging for agents built on **Amazon Bedrock
AgentCore** — LangGraph, Strands Agents, or CrewAI — that emits one
fixed-shape CloudWatch log line your Grafana dashboards can `parse`
reliably, cheaply, and without ever hand-formatting a log string again.

## Why this exists

CloudWatch Logs Insights' `parse` command is a literal, positional string
matcher. It has no notion of "fields" — only fixed text tokens in a fixed
order. A dashboard panel built around:

```
parse message "UserID: * | SessionID: * | TurnID: * | Status: * | InputTokens: * | OutputTokens: * | TTFTMs: * | LatencyMs: *"
```

silently returns **"No data"** the instant any agent's log line drifts
from that exact shape — a swapped field, a quoted number, an extra space.
No error, no warning. Hand-writing that pipe-delimited f-string in every
agent's entrypoint is exactly the kind of thing that breaks quietly and
is expensive to debug after the fact.

This package is the one place that knows the correct shape. Import a
function, call it, and every dashboard reading that log group populates —
whichever framework your agent is written in.

## What you can do with this package

| I want to... | See |
|---|---|
| Log turn metrics automatically, with almost no code change | [Quickstart — tracers](#quickstart--tracers-recommended) |
| Log turn metrics with a single function call, no persistent object | [Quickstart — one-shot wrapper functions](#quickstart--one-shot-wrapper-functions) |
| Log from a framework/loop this package doesn't special-case | [`TurnMetrics`](#anything-else-strands-multi-agent-crewai-flows-a-custom-loop) |
| See which tool/model ran, and its own token usage, per turn | [The `Call metrics` line](#the-call-metrics-line--which-toolagent-ran-and-its-own-token-usage) |
| Get the 4 Grafana dashboards pointed at *my* AWS account/runtime | [Getting the 4 Grafana dashboards](#getting-the-4-grafana-dashboards-for-your-runtime) |
| Generate dashboards from *my own* dashboard JSON, not the bundled 4 | [Bring your own dashboard JSON](#bring-your-own-dashboard-json) |
| Know exactly what gets swapped in my own dashboard JSON, and its limits | [Rules for bring-your-own dashboards](#rules-for-bring-your-own-dashboards) |
| Log my own custom metric line (not turn/call) with a matching Grafana query | [Custom log types — `LogSchema`](#custom-log-types--logschema) |
| Look up one function/class fast | [API reference](#api-reference) |

## Install

```bash
pip install agentcore-dashboard-metrics
```

Zero required dependencies — the framework wrappers are fully duck-typed
against whatever `graph`/`agent`/`crew` object you pass in, so installing
this package never pulls in LangChain, Strands, or CrewAI for you.

## Quickstart — tracers (recommended)

Instantiate a tracer, call `.attach()` once, then use your agent **exactly as
you already do**. No wrapper function, no new return value to unpack — the
log line is emitted as a side effect.

### LangGraph (`create_react_agent`)

```python
from agentcore_dashboard_metrics import LangChainTracer

async def invoke(payload, context):
    ...
    tracer = LangChainTracer(log, user_id=user_id, session_id=session_id)
    tracer.attach(graph)

    async for event in graph.astream_events({"messages": messages}, version="v2"):
        ...  # completely unchanged
```

### Strands Agents

```python
from agentcore_dashboard_metrics import StrandsTracer

async def invoke(payload, context):
    agent = get_or_create_agent()
    tracer = StrandsTracer(log, user_id=user_id, session_id=session_id)
    tracer.attach(agent)

    async for event in agent.stream_async(payload.get("prompt")):
        ...  # completely unchanged
```

### CrewAI

```python
from agentcore_dashboard_metrics import CrewAITracer

async def invoke(payload, context):
    crew = get_or_create_crew()
    tracer = CrewAITracer(log, user_id=user_id, session_id=session_id)
    tracer.attach(crew)

    result = await crew.kickoff_async(inputs={"prompt": payload.get("prompt")})
    # completely unchanged — works for both the non-streaming case and
    # crew.stream=True
```

**How it works:** `.attach()` monkeypatches the exact async streaming method
each framework already exposes for this (`astream_events` / `stream_async` /
`kickoff_async`) so every event/result is yielded through completely
unchanged — verified against a real compiled LangGraph graph, not just a
mock. Re-attaching a fresh tracer to an already-wrapped object (e.g. a
cached graph/agent reused across requests, each with a different
user_id/session_id) swaps which tracer is active instead of stacking
another layer of wrapping.

**The one limitation:** only the wrapped method is traced. `LangChainTracer`
traces `astream_events`, not a separate `ainvoke` call on the same graph;
`StrandsTracer` traces `stream_async`, not the sync `agent(prompt)` call.
Use whichever your entrypoint already calls.

## Quickstart — one-shot wrapper functions

If you don't want a persistent tracer object — e.g. a graph/agent/crew
that's rebuilt fresh on every call anyway — call one of these instead. Same
guarantees, different shape: each does the whole call for you and hands
back `(output_text, turn_id)`.

### LangGraph (`create_react_agent`)

```python
from agentcore_dashboard_metrics import run_agent_turn

async def invoke(payload, context):
    ...
    output, turn_id = await run_agent_turn(
        graph, messages, log,
        user_id=user_id, session_id=session_id,
    )
    return {"result": output, "turn_id": turn_id}
```

### Strands Agents

```python
from agentcore_dashboard_metrics import run_strands_turn

async def invoke(payload, context):
    agent = get_or_create_agent()
    output, turn_id = await run_strands_turn(
        agent, payload.get("prompt"), log,
        user_id=user_id, session_id=session_id,
    )
    return {"result": output, "turn_id": turn_id}
```

### CrewAI

```python
from agentcore_dashboard_metrics import run_crewai_turn

async def invoke(payload, context):
    crew = get_or_create_crew()
    output, turn_id = await run_crewai_turn(
        crew, {"prompt": payload.get("prompt")}, log,
        user_id=user_id, session_id=session_id,
    )
    return {"result": output, "turn_id": turn_id}
```

Each wrapper streams the underlying agent, measures time-to-first-token
and total latency, accumulates token usage across however many model
calls happen in the turn, sets `status` to `success`/`error`, and emits
the log line — all of it, so your entrypoint has nothing left to
hand-format.

### Anything else (Strands multi-agent, CrewAI Flows, a custom loop)

```python
from agentcore_dashboard_metrics import TurnMetrics

with TurnMetrics(log, user_id=user_id, session_id=session_id) as m:
    result = my_own_agent_call(...)
    m.add_usage(input_tokens=result.input_tokens,
                output_tokens=result.output_tokens)
    m.mark_first_token()  # optional, only if you stream
# the correctly-shaped line is logged automatically on exit,
# including on an exception (status is set to "error" for you,
# and the exception is re-raised — never swallowed)
```

### Already computed everything yourself?

```python
from agentcore_dashboard_metrics import log_turn_metrics

log_turn_metrics(
    log, user_id=user_id, session_id=session_id, turn_id=turn_id,
    status="success", input_tokens=412, output_tokens=88,
    ttft_ms=640.12, latency_ms=1820.55,
)
```

## The log line

Every path above produces exactly this shape via a single `log.info(...)`
call:

```
Published metrics — UserID: nakul | SessionID: 826fc1dc-... | TurnID: 3f9c2e1a-... | Status: success | InputTokens: 412 | OutputTokens: 88 | TTFTMs: 640.12 | LatencyMs: 1820.55
```

Build your CloudWatch Logs Insights `parse` pattern against exactly that
text and every dashboard panel — `stats sum(input_tokens) by user_id`,
`stats avg(ttft_ms)`, `stats count(*) by status`, etc. — works against any
agent using this package, regardless of which of the three frameworks it's
built on.

### Field reference

| Field | Type | Notes |
|---|---|---|
| `UserID` | string | End user identifier |
| `SessionID` | string | Conversation/session identifier |
| `TurnID` | string | Unique per invocation; auto-generated (`new_turn_id()`) if not supplied |
| `Status` | string | Single word, lowercased, no spaces/pipes — `success` or `error` |
| `InputTokens` / `OutputTokens` | int | Summed across every model call in the turn |
| `TTFTMs` | float | Time to first streamed token, ms; `-1.00` if the call shape has no token-level stream (e.g. CrewAI's non-streaming `kickoff_async`) |
| `LatencyMs` | float | Total wall-clock time for the turn |

All fields are sanitized defensively — `None`, empty strings, and stray
`|`/newline characters in ids or status are coerced into something safe
rather than corrupting the `parse` pattern.

### Adding a new field

If you need an additional metric (e.g. `ModelID`, `ToolUsed`), append it
**after `LatencyMs`** in your own logging so it doesn't shift the position
of any field existing dashboards already parse — `parse` ignores trailing
fields a given query doesn't ask for.

## The `Call metrics` line — which tool/agent ran, and its own token usage

`Published metrics` gives you the turn as a whole. It can't answer "which
tools or sub-agents actually ran during this turn, and how many tokens did
each one use" — that needs a second, per-call line, correlated back to the
turn via `TurnID`:

```
Call metrics — UserID: nakul | SessionID: sess-1 | TurnID: turn-1 | CallIndex: 0 | Type: model | Name: anthropic.claude-3-sonnet | Status: success | InputTokens: 412 | OutputTokens: 88
Call metrics — UserID: nakul | SessionID: sess-1 | TurnID: turn-1 | CallIndex: 1 | Type: tool | Name: get_weather | Status: success | InputTokens: 0 | OutputTokens: 0
```

One line per LLM step / tool call / sub-agent step, in order (`CallIndex`
starting at 0), kept as its own log statement rather than a nested array
inside the turn line — CloudWatch Logs Insights can't unnest/query a nested
array per-row, but it can group and aggregate flat lines like these by
`name`.

**You get this for free with LangGraph and Strands** — nothing to change in
your own code.

- **LangGraph:** `run_agent_turn` and `LangChainTracer` record one `Call
  metrics` entry per `on_chat_model_end` (type `model`) and per
  `on_tool_end`/`on_tool_error` (type `tool`, status `success`/`error`).
- **Strands:** `run_strands_turn` and `StrandsTracer` record one entry per
  model cycle (type `model`, that cycle's own tokens) and one per tool call
  (type `tool`), read from the agent's own metrics. Counts are for *this
  turn only*, even when one agent instance serves a whole session.
- **CrewAI:** not wired up yet — see `add_call()` below to do it yourself.

Without `Call metrics` lines the three "Tools & Agents" panels on the turn
dashboard stay empty.

| Field | Type | Notes |
|---|---|---|
| `CallIndex` | int | 0-based order within the turn |
| `Type` | string | e.g. `model` or `tool` |
| `Name` | string | Tool name, or the model id for an LLM step |
| `Status` | string | `success` or `error` |
| `InputTokens` / `OutputTokens` | int | That call's own usage — usually `0` for a tool call |

To record calls yourself (any framework, via `TurnMetrics` or inside a
`with` block):

```python
with TurnMetrics(log, user_id=user_id, session_id=session_id) as m:
    ...
    m.add_call(name="get_weather", type="tool")
    m.add_call(name=model_id, type="model", input_tokens=412, output_tokens=88)
# one `Call metrics` line per add_call(), logged automatically on exit,
# right after the turn's `Published metrics` line
```

`add_call()` is independent of `add_usage()` — recording a call's own
token usage does **not** also add it to the turn-level total; call both if
you want each.

## Getting the 4 Grafana dashboards for *your* runtime

This package's 4 dashboards (home / session / user / turn based) ship as
templates — their queries have an AWS account id, region, and log group
baked in, so they need to be pointed at *your* runtime before they'll show
your data. `pip install` gives you a real terminal command for that — no
`python -m`, no script to write:

```bash
agentcore-dashboard-metrics create --arn arn:aws:bedrock-agentcore:us-east-1:111122223333:runtime/MyAgent-ab12cd34ef
```

```
Parsing runtime ARN
  region      us-east-1
  account id  111122223333
  runtime id  MyAgent-ab12cd34ef

Generating dashboards
  [1/4] ███████░░░░░░░░░░░░░░░░░░░░░  25%  home-dashboard.json
  [2/4] ██████████████░░░░░░░░░░░░░░  50%  session-based-dashbaord.json
  [3/4] █████████████████████░░░░░░░  75%  user-based-dashboard.json
  [4/4] ████████████████████████████ 100%  turn-based-dashboard.json

✔ 4 dashboards created successfully
→ /home/you/dashboards
```

Account id, region, and runtime id are all parsed out of that one ARN (the
same one `agentcore deploy` prints and the AgentCore console shows) — no
AWS credentials are read or used, this only edits JSON on disk. Import the
4 files it created into Grafana and you're done.

```bash
agentcore-dashboard-metrics create \
  --arn arn:aws:bedrock-agentcore:us-east-1:111122223333:runtime/MyAgent-ab12cd34ef \
  --output my-dashboards \                      # default: ./dashboards
  --log-group /my/custom/log/group \            # only if you don't use AgentCore's
                                                 # default naming convention
  --uids '{"session": "myses123"}'              # only if the default uids collide
                                                 # with something already in your Grafana

agentcore-dashboard-metrics create --arn ... --quiet     # prints only the final path — for scripting
```

### Bring your own dashboard JSON

Not limited to this package's 4 bundled dashboards. Export any dashboard
from Grafana as JSON, put the file(s) in a folder, and point them at your
runtime — **no editing, no placeholders**:

```bash
agentcore-dashboard-metrics create --arn <your-runtime-arn> --template-dir my-dashboards/
```

The package finds the account id, region, runtime id and log group the
dashboard *currently* points at (from the ARNs, `accountId`/`region` keys
and log-group names already in the JSON) and swaps them for the ones in
`--arn`. Dashboard uids and everything else are left exactly as they are.

### Rules for bring-your-own dashboards

What gets swapped, and where it can go wrong:

1. **Must be valid JSON**, as Grafana exports it. Parsed with `json.loads()`
   first; a malformed file fails immediately.
2. **The old values are found in these places:**
   - CloudWatch log-group ARNs — `arn:aws:logs:<region>:<account>:log-group:<name>:*`
   - runtime ARNs — `arn:aws:bedrock-agentcore:<region>:<account>:runtime/<id>`
   - AgentCore log group names — `/aws/bedrock-agentcore/runtimes/<id>-DEFAULT`
   - `accountId` and `region` keys, and `logGroups[].name`

   Every occurrence of the old account id, region, runtime id and log group
   is replaced, in string values only (JSON keys are never changed).
3. **Other log groups are left alone.** A dashboard that also queries e.g.
   `/aws/application-signals/data` keeps that untouched — only the
   runtime's own log group is swapped. (Its account id and region are still
   swapped, since they're part of the same ARN.)
4. **Several runtimes in one file are all replaced.** If a file mentions
   more than one runtime, account or region, every one of them is replaced
   by the values from `--arn`.
5. **Nothing to replace is an error** — a file with no ARN, `accountId`
   or `region` values raises `ValueError` instead of silently writing an
   unchanged copy.
6. **Uids stay as they are** unless you pass `--uids '{"old_uid": "new_uid"}'`
   (every occurrence — the dashboard's own uid and links pointing at it).
   For the bundled dashboards you can use the names
   `home`/`session`/`user`/`turn` instead of their uids.
7. **Filenames must be unique within one `--template-dir`.** The output
   filename is the input filename, and only `*.json` files at the top
   level of the folder are read (not recursively).
8. **`--log-group NAME`** sets the new log group when yours doesn't follow
   AgentCore's `/aws/bedrock-agentcore/runtimes/<runtime_id>-DEFAULT` naming.

Prefer calling it from Python instead of the shell? Same thing, one function:

```python
from agentcore_dashboard_metrics.dashboards import generate_dashboards

written = generate_dashboards(
    runtime_arn="arn:aws:bedrock-agentcore:us-east-1:111122223333:runtime/MyAgent-ab12cd34ef",
    # output_dir="dashboards", log_group_name=..., dashboard_uids=... — same optional overrides
    # progress_callback=lambda index, total, filename, out_path: ...,  # what the CLI uses internally
)
```

Raises `RuntimeArnError` if `runtime_arn` isn't a well-formed Bedrock
AgentCore runtime ARN (`arn:aws:bedrock-agentcore:<region>:<account_id>:runtime/<runtime_id>`).

For your own dashboard JSON(s), use `generate_from_templates()` instead —
`generate_dashboards()` is a thin wrapper over it for the 4 bundled files:

```python
from agentcore_dashboard_metrics.dashboards import generate_from_templates

written = generate_from_templates(
    template_paths=["my-dashboard.json"],
    runtime_arn="arn:aws:bedrock-agentcore:us-east-1:111122223333:runtime/MyAgent-ab12cd34ef",
)
```

Changing uids as well:

```python
written = generate_from_templates(
    template_paths=["my-dashboard.json"],
    runtime_arn="arn:aws:bedrock-agentcore:us-east-1:111122223333:runtime/MyAgent-ab12cd34ef",
    dashboard_uids={"old_uid": "new_uid"},   # optional — uids are otherwise untouched
)
```

To see what would be replaced without converting anything, use
`detect_source_values(json.load(open("my-dashboard.json")))` — it returns
the account ids, regions, runtime ids and log groups it found.

## Custom log types — `LogSchema`

`Published metrics` and `Call metrics` cover turn/call-level logging, but
you can register your own pipe-delimited log shape the same way — one
place that knows the field order, used to both format the line and
generate the matching CloudWatch Logs Insights `parse` pattern, so the two
can't silently drift apart:

```python
from agentcore_dashboard_metrics import LogSchema, register_log_schema, log_custom_metrics

cache_schema = LogSchema("Cache metrics", [
    ("user_id", "UserID"),
    ("cache_key", "CacheKey"),
    ("hit", "Hit"),
])
register_log_schema(cache_schema)

log_custom_metrics(log, "Cache metrics", user_id=user_id, cache_key="prompt-v3", hit=True)
# -> "Cache metrics — UserID: nakul | CacheKey: prompt-v3 | Hit: True"

cache_schema.parse_pattern
# -> 'parse @message "Cache metrics — UserID: * | CacheKey: * | Hit: *" as user_id, cache_key, hit'
```

Paste `.parse_pattern` straight into a new Grafana panel's CloudWatch Logs
Insights query — no hand-written query to keep in sync with the log line.

## API reference

| Name | Use for |
|---|---|
| `LangChainTracer(log, *, user_id, session_id).attach(graph)` | LangGraph `create_react_agent` — instruments `astream_events` |
| `StrandsTracer(log, *, user_id, session_id).attach(agent)` | Strands `Agent` — instruments `stream_async` |
| `CrewAITracer(log, *, user_id, session_id).attach(crew)` | CrewAI `Crew` — instruments `kickoff_async` |
| `run_agent_turn(graph, messages, log, *, user_id, session_id, turn_id=None)` | LangGraph, one-shot |
| `run_strands_turn(agent, prompt, log, *, user_id, session_id, turn_id=None)` | Strands, one-shot |
| `run_crewai_turn(crew, inputs, log, *, user_id, session_id, turn_id=None)` | CrewAI, one-shot |
| `TurnMetrics(log, *, user_id, session_id, turn_id=None)` | Context manager for any other framework/loop |
| `log_turn_metrics(log, *, user_id, session_id, turn_id, status, input_tokens, output_tokens, ttft_ms, latency_ms)` | Raw formatter, if you've already computed everything |
| `log_call_metrics(log, *, user_id, session_id, turn_id, calls)` | Raw formatter for the per-call breakdown — `calls` is a list of `{name, type, status, input_tokens, output_tokens}` dicts |
| `new_turn_id()` | Generates a fresh turn id |
| `LogSchema(name, fields)` / `register_log_schema(schema)` / `log_custom_metrics(log, name, **fields)` | Define and emit your own pipe-delimited log line + matching `parse` pattern |
| `agentcore_dashboard_metrics.dashboards.generate_dashboards(*, runtime_arn, output_dir="dashboards", log_group_name=None, dashboard_uids=None)` | Writes the 4 bundled Grafana dashboards, parameterized for your runtime ARN |
| `agentcore_dashboard_metrics.dashboards.generate_from_templates(*, template_paths, runtime_arn, output_dir="dashboards", log_group_name=None, dashboard_uids=None)` | Same, for your own dashboard JSON file(s) — old values are detected from the JSON and swapped for `runtime_arn`'s |
| `agentcore_dashboard_metrics.dashboards.render_from_templates(*, template_paths, runtime_arn, log_group_name=None, dashboard_uids=None)` | Like `generate_from_templates()` but returns `{filename: rendered_text}` without writing |
| `agentcore_dashboard_metrics.dashboards.detect_source_values(dashboard_json)` | The account ids, regions, runtime ids and log groups found in a parsed dashboard — what would be swapped |
| `agentcore_dashboard_metrics.dashboards.default_template_paths()` | The 4 bundled template file paths, in generation order |
| `agentcore_dashboard_metrics.dashboards.parse_runtime_arn(runtime_arn)` | Parses a runtime ARN into `{region, account_id, runtime_id}` |

All three `run_*_turn()` functions are `async def` and return
`(output_text, turn_id)`. On an exception from the underlying
graph/agent/crew, the metrics line is still logged with
`status="error"` and the exception is re-raised unchanged — the same
guarantee applies to the tracers' wrapped methods.

## License

MIT
