Metadata-Version: 2.5
Name: agentcore-dashboard-metrics
Version: 0.1.0
Summary: Plug-and-play turn-metrics logging for LangGraph/Strands/CrewAI agents on Bedrock AgentCore, shaped for CloudWatch Logs Insights + Grafana dashboards.
Author-email: nakulan <nakult721@gmail.com>
License-Expression: MIT
License-File: LICENSE
Keywords: agentcore,bedrock,cloudwatch,crewai,grafana,langchain,langgraph,logging,observability,strands-agents
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: System :: Logging
Classifier: Topic :: System :: Monitoring
Classifier: Typing :: Typed
Requires-Python: >=3.9
Provides-Extra: dev
Requires-Dist: build; extra == 'dev'
Requires-Dist: pytest; extra == 'dev'
Requires-Dist: twine; extra == 'dev'
Description-Content-Type: text/markdown

# agentcore-dashboard-metrics

Plug-and-play turn-metrics logging for agents built on **Amazon Bedrock
AgentCore** — LangGraph, Strands Agents, or CrewAI — that emits one
fixed-shape CloudWatch log line your Grafana dashboards can `parse`
reliably, cheaply, and without ever hand-formatting a log string again.

## Why this exists

CloudWatch Logs Insights' `parse` command is a literal, positional string
matcher. It has no notion of "fields" — only fixed text tokens in a fixed
order. A dashboard panel built around:

```
parse message "UserID: * | SessionID: * | TurnID: * | Status: * | InputTokens: * | OutputTokens: * | TTFTMs: * | LatencyMs: *"
```

silently returns **"No data"** the instant any agent's log line drifts
from that exact shape — a swapped field, a quoted number, an extra space.
No error, no warning. Hand-writing that pipe-delimited f-string in every
agent's entrypoint is exactly the kind of thing that breaks quietly and
is expensive to debug after the fact.

This package is the one place that knows the correct shape. Import a
function, call it, and every dashboard reading that log group populates —
whichever framework your agent is written in.

## Install

```bash
pip install agentcore-dashboard-metrics
```

Zero required dependencies — the framework wrappers are fully duck-typed
against whatever `graph`/`agent`/`crew` object you pass in, so installing
this package never pulls in LangChain, Strands, or CrewAI for you.

## Quickstart — tracers (recommended)

Instantiate a tracer, call `.attach()` once, then use your agent **exactly as
you already do**. No wrapper function, no new return value to unpack — the
log line is emitted as a side effect.

### LangGraph (`create_react_agent`)

```python
from agentcore_dashboard_metrics import LangChainTracer

async def invoke(payload, context):
    ...
    tracer = LangChainTracer(log, user_id=user_id, session_id=session_id)
    tracer.attach(graph)

    async for event in graph.astream_events({"messages": messages}, version="v2"):
        ...  # completely unchanged
```

### Strands Agents

```python
from agentcore_dashboard_metrics import StrandsTracer

async def invoke(payload, context):
    agent = get_or_create_agent()
    tracer = StrandsTracer(log, user_id=user_id, session_id=session_id)
    tracer.attach(agent)

    async for event in agent.stream_async(payload.get("prompt")):
        ...  # completely unchanged
```

### CrewAI

```python
from agentcore_dashboard_metrics import CrewAITracer

async def invoke(payload, context):
    crew = get_or_create_crew()
    tracer = CrewAITracer(log, user_id=user_id, session_id=session_id)
    tracer.attach(crew)

    result = await crew.kickoff_async(inputs={"prompt": payload.get("prompt")})
    # completely unchanged — works for both the non-streaming case and
    # crew.stream=True
```

**How it works:** `.attach()` monkeypatches the exact async streaming method
each framework already exposes for this (`astream_events` / `stream_async` /
`kickoff_async`) so every event/result is yielded through completely
unchanged — verified against a real compiled LangGraph graph, not just a
mock. Re-attaching a fresh tracer to an already-wrapped object (e.g. a
cached graph/agent reused across requests, each with a different
user_id/session_id) swaps which tracer is active instead of stacking
another layer of wrapping.

**The one limitation:** only the wrapped method is traced. `LangChainTracer`
traces `astream_events`, not a separate `ainvoke` call on the same graph;
`StrandsTracer` traces `stream_async`, not the sync `agent(prompt)` call.
Use whichever your entrypoint already calls.

## Quickstart — one-shot wrapper functions

If you don't want a persistent tracer object — e.g. a graph/agent/crew
that's rebuilt fresh on every call anyway — call one of these instead. Same
guarantees, different shape: each does the whole call for you and hands
back `(output_text, turn_id)`.

### LangGraph (`create_react_agent`)

```python
from agentcore_dashboard_metrics import run_agent_turn

async def invoke(payload, context):
    ...
    output, turn_id = await run_agent_turn(
        graph, messages, log,
        user_id=user_id, session_id=session_id,
    )
    return {"result": output, "turn_id": turn_id}
```

### Strands Agents

```python
from agentcore_dashboard_metrics import run_strands_turn

async def invoke(payload, context):
    agent = get_or_create_agent()
    output, turn_id = await run_strands_turn(
        agent, payload.get("prompt"), log,
        user_id=user_id, session_id=session_id,
    )
    return {"result": output, "turn_id": turn_id}
```

### CrewAI

```python
from agentcore_dashboard_metrics import run_crewai_turn

async def invoke(payload, context):
    crew = get_or_create_crew()
    output, turn_id = await run_crewai_turn(
        crew, {"prompt": payload.get("prompt")}, log,
        user_id=user_id, session_id=session_id,
    )
    return {"result": output, "turn_id": turn_id}
```

Each wrapper streams the underlying agent, measures time-to-first-token
and total latency, accumulates token usage across however many model
calls happen in the turn, sets `status` to `success`/`error`, and emits
the log line — all of it, so your entrypoint has nothing left to
hand-format.

### Anything else (Strands multi-agent, CrewAI Flows, a custom loop)

```python
from agentcore_dashboard_metrics import TurnMetrics

with TurnMetrics(log, user_id=user_id, session_id=session_id) as m:
    result = my_own_agent_call(...)
    m.add_usage(input_tokens=result.input_tokens,
                output_tokens=result.output_tokens)
    m.mark_first_token()  # optional, only if you stream
# the correctly-shaped line is logged automatically on exit,
# including on an exception (status is set to "error" for you,
# and the exception is re-raised — never swallowed)
```

### Already computed everything yourself?

```python
from agentcore_dashboard_metrics import log_turn_metrics

log_turn_metrics(
    log, user_id=user_id, session_id=session_id, turn_id=turn_id,
    status="success", input_tokens=412, output_tokens=88,
    ttft_ms=640.12, latency_ms=1820.55,
)
```

## The log line

Every path above produces exactly this shape via a single `log.info(...)`
call:

```
Published metrics — UserID: nakul | SessionID: 826fc1dc-... | TurnID: 3f9c2e1a-... | Status: success | InputTokens: 412 | OutputTokens: 88 | TTFTMs: 640.12 | LatencyMs: 1820.55
```

Build your CloudWatch Logs Insights `parse` pattern against exactly that
text and every dashboard panel — `stats sum(input_tokens) by user_id`,
`stats avg(ttft_ms)`, `stats count(*) by status`, etc. — works against any
agent using this package, regardless of which of the three frameworks it's
built on.

### Field reference

| Field | Type | Notes |
|---|---|---|
| `UserID` | string | End user identifier |
| `SessionID` | string | Conversation/session identifier |
| `TurnID` | string | Unique per invocation; auto-generated (`new_turn_id()`) if not supplied |
| `Status` | string | Single word, lowercased, no spaces/pipes — `success` or `error` |
| `InputTokens` / `OutputTokens` | int | Summed across every model call in the turn |
| `TTFTMs` | float | Time to first streamed token, ms; `-1.00` if the call shape has no token-level stream (e.g. CrewAI's non-streaming `kickoff_async`) |
| `LatencyMs` | float | Total wall-clock time for the turn |

All fields are sanitized defensively — `None`, empty strings, and stray
`|`/newline characters in ids or status are coerced into something safe
rather than corrupting the `parse` pattern.

### Adding a new field

If you need an additional metric (e.g. `ModelID`, `ToolUsed`), append it
**after `LatencyMs`** in your own logging so it doesn't shift the position
of any field existing dashboards already parse — `parse` ignores trailing
fields a given query doesn't ask for.

## API reference

| Name | Use for |
|---|---|
| `LangChainTracer(log, *, user_id, session_id).attach(graph)` | LangGraph `create_react_agent` — instruments `astream_events` |
| `StrandsTracer(log, *, user_id, session_id).attach(agent)` | Strands `Agent` — instruments `stream_async` |
| `CrewAITracer(log, *, user_id, session_id).attach(crew)` | CrewAI `Crew` — instruments `kickoff_async` |
| `run_agent_turn(graph, messages, log, *, user_id, session_id, turn_id=None)` | LangGraph, one-shot |
| `run_strands_turn(agent, prompt, log, *, user_id, session_id, turn_id=None)` | Strands, one-shot |
| `run_crewai_turn(crew, inputs, log, *, user_id, session_id, turn_id=None)` | CrewAI, one-shot |
| `TurnMetrics(log, *, user_id, session_id, turn_id=None)` | Context manager for any other framework/loop |
| `log_turn_metrics(log, *, user_id, session_id, turn_id, status, input_tokens, output_tokens, ttft_ms, latency_ms)` | Raw formatter, if you've already computed everything |
| `new_turn_id()` | Generates a fresh turn id |

All three `run_*_turn()` functions are `async def` and return
`(output_text, turn_id)`. On an exception from the underlying
graph/agent/crew, the metrics line is still logged with
`status="error"` and the exception is re-raised unchanged — the same
guarantee applies to the tracers' wrapped methods.

## License

MIT
