◆ Developer guide

Stop hand-formatting log lines

agentcore-dashboard-metrics is a small Python package that logs exactly the metrics line your Grafana dashboards need — for agents on Bedrock AgentCore built with LangGraph, Strands, or CrewAI. It also knows which tool/model ran inside each turn, and generates all 4 dashboards parameterized for your own AWS account.

The problem it solves

CloudWatch Logs Insights' parse command is a literal, positional string matcher — it has no notion of a "field," only fixed text tokens in a fixed order. A dashboard panel built to read a specific shape returns a silent "No data" the moment any agent's log line drifts from it: a swapped field, a quoted number, one extra space. No error, no warning — just an empty chart.

This is the exact string every dashboard panel expects, and the one this package emits for you on every turn:

Published metrics — UserID: nakul | SessionID: 826fc1dc-… | TurnID: 3f9c2e1a-… | Status: success | InputTokens: 412 | OutputTokens: 88 | TTFTMs: 640.12 | LatencyMs: 1820.55
UserID / SessionID
Who, and which conversation. Drive Home & the per-user/session dashboards.
TurnID
One value per invocation. Auto-generated if you don't supply one.
Status
Single word, no spaces — success or error.
InputTokens / OutputTokens
Summed across every model call made during the turn.
TTFTMs / LatencyMs
Time-to-first-token and total wall-clock time, in milliseconds.
Getting started

Install

Zero required runtime dependencies — the framework wrappers are fully duck-typed against whatever graph / agent / crew object you pass them, so installing this package never pulls in LangChain, Strands, or CrewAI on your behalf.

pip install agentcore-dashboard-metrics
Plug and play · recommended

Instantiate a tracer, .attach() it, done

No wrapper function, no new return value to unpack. Use your agent exactly as you already do — the log line is emitted as a side effect.

from agentcore_dashboard_metrics import LangChainTracer

async def invoke(payload, context):
    # ...build graph, messages...
    tracer = LangChainTracer(log, user_id=user_id, session_id=session_id)
    tracer.attach(graph)

    # unchanged from here — tracing happens as a side effect:
    async for event in graph.astream_events({"messages": messages}, version="v2"):
        ...

How it works: .attach() monkeypatches the exact async streaming method each framework already exposes (astream_events / stream_async / kickoff_async), yielding every event/result through completely unchanged — verified against a real compiled LangGraph graph, not just a mock. Re-attaching a fresh tracer to an already-wrapped object (a cached graph/agent reused across requests) swaps the active tracer instead of stacking another layer of wrapping.

One limitation: only the wrapped method is traced — LangChainTracer traces astream_events, not a separate ainvoke call on the same graph; StrandsTracer traces stream_async, not the sync agent(prompt) call.

No persistent tracer object needed

One-shot functions

If your graph/agent/crew is rebuilt fresh on every call anyway, skip the tracer and just call one of these — same guarantees, hands back the result directly.

from agentcore_dashboard_metrics import run_agent_turn

async def invoke(payload, context):
    # ...build graph, messages...
    output, turn_id = await run_agent_turn(
        graph, messages, log,
        user_id=user_id, session_id=session_id,
    )
    return {"result": output, "turn_id": turn_id}

What happens under the hood: streams the agent, marks the first output chunk for TTFT, sums token usage across however many model calls the turn makes, records which tool/model ran (see Call metrics below), sets status to success or error, and emits the log line — on an exception, the line is still logged with status="error" and the exception is re-raised unchanged.

Anything else

Strands multi-agent, CrewAI Flows, or a custom loop

Drop to the framework-agnostic context manager. It handles timing and the log line; you just report usage as you get it.

from agentcore_dashboard_metrics import TurnMetrics

with TurnMetrics(log, user_id=user_id, session_id=session_id) as m:
    result = my_own_agent_call(...)
    m.add_usage(input_tokens=result.input_tokens,
                output_tokens=result.output_tokens)
    m.mark_first_token()  # optional — only if you stream
# the log line is emitted automatically on exit, including on
# an exception (status is forced to "error"; nothing is swallowed)

Already computed everything yourself? Skip straight to the formatter:

from agentcore_dashboard_metrics import log_turn_metrics

log_turn_metrics(
    log, user_id=user_id, session_id=session_id, turn_id=turn_id,
    status="success", input_tokens=412, output_tokens=88,
    ttft_ms=640.12, latency_ms=1820.55,
)
The contract

Rules that keep every panel populated

The package enforces these for you — sanitizing bad input rather than corrupting the line — but it's worth knowing what it's protecting against if you ever log this by hand.

  • 01 Line starts with Published metrics — (em dash, not a hyphen) — every panel filters on it.
  • 02 Fields stay in order: UserID → SessionID → TurnID → Status → InputTokens → OutputTokens → TTFTMs → LatencyMs.
  • 03 Separator is exactly " | " — space, pipe, space.
  • 04 Numbers are bare and unquoted — stats sum()/avg() need real numerics.
  • 05 Status is one lowercase word, no spaces or pipes.
  • 06 Exactly one line per turn — double-logging breaks every count() panel.
  • 07 IDs are always non-empty, stable strings — they double as click-through link values.

TTFT is honest, not universal. ttft_ms is only meaningful when the framework streams individual chunks. CrewAI's common non-streaming kickoff_async call has no token-level stream to time, so run_crewai_turn reports -1 there — the same convention used for any turn with no streamed output.

Which tool or agent ran

The Call metrics line

Published metrics gives you the turn as a whole — it can't say which tool or model ran, or that call's own token usage. A second, correlated line answers that:

Call metrics — UserID: nakul | SessionID: sess-1 | TurnID: turn-1 | CallIndex: 0 | Type: model | Name: claude-3-sonnet | Status: success | InputTokens: 412 | OutputTokens: 88

One line per LLM step / tool call, in order (CallIndex starting at 0), correlated back to the turn via TurnID — kept separate rather than nested inside the turn line, since CloudWatch Logs Insights can't unnest an array per-row, but it can group flat lines like these by name.

You get this for free with LangGraph and Strands — nothing to change in your own code. LangGraph: run_agent_turn and LangChainTracer record one entry per model step and per tool call. Strands: run_strands_turn and StrandsTracer record one entry per model cycle and per tool call, for that turn only, even when one agent serves a whole session. CrewAI isn't wired up yet; use add_call() below to do it yourself. Without these lines the three "Tools & Agents" panels on the turn dashboard stay empty.

with TurnMetrics(log, user_id=user_id, session_id=session_id) as m:
    ...
    m.add_call(name="get_weather", type="tool")
    m.add_call(name=model_id, type="model", input_tokens=412, output_tokens=88)
# one `Call metrics` line per add_call(), logged automatically on exit,
# right after the turn's `Published metrics` line

add_call() is independent of add_usage() — recording a call's own tokens does not also add them to the turn-level total; call both if you want each.

On the Grafana side

Four dashboards, one log group

Once your agent logs this shape, these dashboards read it directly — no code changes, no per-agent wiring.

OVERVIEW

Home

Top sessions/users by token usage, system error rate, latency at a glance. Start here.

$SessionID

Session-based

Everything that happened in one conversation — invocation rate, TTFT, success ratio, owner.

$UserID

User-based

One user's usage across all sessions — token spend, invocation trend, session list.

$TurnID

Turn-based

The exact detail of one exchange — latency, TTFT, tokens, status, and which tools/models ran and their own token usage.

Click-through navigation

Every ID rendered as a link jumps to the related dashboard with that value pre-filled and the time range carried over:

Home→user_id User-based→session_id Session-based→turn_id Turn-based

Getting these dashboards for your runtime

The 4 dashboards were built against one deployment, so their queries carry that deployment's AWS account id, region, runtime id and log group. One command points all 4 at yours. Everything comes from the runtime ARN you already have; no AWS credentials are read, the command only edits JSON on disk.

agentcore-dashboard-metrics create \
  --arn arn:aws:bedrock-agentcore:us-east-1:111122223333:runtime/MyAgent-ab12cd34ef
# writes ./dashboards/ with all 4 JSON files, ready to import into Grafana

Or from Python: generate_dashboards(runtime_arn=...) from agentcore_dashboard_metrics.dashboards.

Your own dashboard JSON

Export any dashboard from Grafana and point it at your runtime. There are no placeholders to add: the package finds the account id, region, runtime id and log group the file currently points at and swaps them for the ones in --arn. Dashboard uids and everything else stay as they are.

agentcore-dashboard-metrics create --arn <your-runtime-arn> --template-dir my-dashboards/ --output out
  • Swapped: log-group ARNs, runtime ARNs, /aws/bedrock-agentcore/runtimes/<id>-DEFAULT names, and accountId / region fields. Only string values change, never keys.
  • Other log groups, such as /aws/application-signals/data, are left alone.
  • A file that mentions several runtimes, accounts or regions has all of them replaced. A file with nothing to replace is an error.
  • --log-group sets the new log group when yours isn't AgentCore's default name.
  • Uids change only when you pass --uids '{"old_uid": "new_uid"}'. For the bundled dashboards the names home, session, user, turn work.
Your own log lines

Custom log types

Register a pipe-delimited line of your own and get the matching CloudWatch Logs Insights parse pattern from the same definition, so the log line and the panel query can't drift apart.

from agentcore_dashboard_metrics import LogSchema, register_log_schema, log_custom_metrics

schema = register_log_schema(LogSchema("Cache metrics", [
    ("user_id", "UserID"), ("cache_key", "CacheKey"), ("hit", "Hit"),
]))
log_custom_metrics(log, "Cache metrics", user_id="nakul", cache_key="prompt-v3", hit=True)
# Cache metrics — UserID: nakul | CacheKey: prompt-v3 | Hit: True

schema.parse_pattern
# parse @message "Cache metrics — UserID: * | CacheKey: * | Hit: *" as user_id, cache_key, hit
Reference

API

NameUse for
LangChainTracer(log, *, user_id, session_id).attach(graph)LangGraph — instruments astream_events
StrandsTracer(log, *, user_id, session_id).attach(agent)Strands — instruments stream_async
CrewAITracer(log, *, user_id, session_id).attach(crew)CrewAI — instruments kickoff_async
run_agent_turn(graph, messages, log, *, user_id, session_id, turn_id=None)LangGraph, one-shot
run_strands_turn(agent, prompt, log, *, user_id, session_id, turn_id=None)Strands, one-shot
run_crewai_turn(crew, inputs, log, *, user_id, session_id, turn_id=None)CrewAI, one-shot
TurnMetrics(log, *, user_id, session_id, turn_id=None)Context manager for anything else
log_turn_metrics(log, *, user_id, session_id, turn_id, status, input_tokens, output_tokens, ttft_ms, latency_ms)Raw formatter — turn line
log_call_metrics(log, *, user_id, session_id, turn_id, calls)Raw formatter — per-call breakdown
new_turn_id()Generates a fresh turn id
LogSchema(name, fields) / register_log_schema() / log_custom_metrics(log, name, **fields)Your own log line plus its parse pattern
dashboards.generate_dashboards(*, runtime_arn, output_dir="dashboards", ...)Writes the 4 dashboards for your runtime
dashboards.generate_from_templates(*, template_paths, runtime_arn, ...)Same, for your own dashboard JSON files
dashboards.detect_source_values(dashboard_json)The account ids, regions, runtime ids and log groups that would be swapped
dashboards.parse_runtime_arn(runtime_arn)Parses an ARN into {region, account_id, runtime_id}

All three run_*_turn() functions are async def and return (output_text, turn_id). On an exception, the metrics line is still logged with status="error" and the exception re-raised — the same guarantee applies to the tracers' wrapped methods.

When a panel is empty

"No data" checklist

  1. Is your agent actually logging? Confirm Published metrics lines are reaching CloudWatch at all.
  2. Does the line match the contract in this doc? Copy one real line out of CloudWatch and run it through the dashboard's parse pattern manually.
  3. Is the time range wide enough? Dashboards default to now-24h → now.
  4. Is the variable value exact? $SessionID/$UserID/$TurnID do exact string matches.