Stop hand-formatting log lines
agentcore-dashboard-metrics is a small Python package that
logs exactly the metrics line your Grafana dashboards need — for agents on
Bedrock AgentCore built with LangGraph, Strands, or CrewAI. It also
knows which tool/model ran inside each turn, and generates all 4 dashboards
parameterized for your own AWS account.
The problem it solves
CloudWatch Logs Insights' parse command is a literal,
positional string matcher — it has no notion of a "field," only fixed text tokens in a
fixed order. A dashboard panel built to read a specific shape returns a silent
"No data" the moment any agent's log line drifts from it: a swapped
field, a quoted number, one extra space. No error, no warning — just an empty chart.
This is the exact string every dashboard panel expects, and the one this package emits for you on every turn:
success or error.Install
Zero required runtime dependencies — the framework wrappers are fully duck-typed
against whatever graph / agent /
crew object you pass them, so installing this package never
pulls in LangChain, Strands, or CrewAI on your behalf.
Instantiate a tracer, .attach() it, done
No wrapper function, no new return value to unpack. Use your agent exactly as you already do — the log line is emitted as a side effect.
from agentcore_dashboard_metrics import LangChainTracer async def invoke(payload, context): # ...build graph, messages... tracer = LangChainTracer(log, user_id=user_id, session_id=session_id) tracer.attach(graph) # unchanged from here — tracing happens as a side effect: async for event in graph.astream_events({"messages": messages}, version="v2"): ...
from agentcore_dashboard_metrics import StrandsTracer async def invoke(payload, context): agent = get_or_create_agent() tracer = StrandsTracer(log, user_id=user_id, session_id=session_id) tracer.attach(agent) # unchanged from here: async for event in agent.stream_async(payload.get("prompt")): ...
from agentcore_dashboard_metrics import CrewAITracer async def invoke(payload, context): crew = get_or_create_crew() tracer = CrewAITracer(log, user_id=user_id, session_id=session_id) tracer.attach(crew) # unchanged — works for non-streaming AND crew.stream=True: result = await crew.kickoff_async(inputs={"prompt": payload.get("prompt")})
How it works: .attach() monkeypatches
the exact async streaming method each framework already exposes
(astream_events / stream_async /
kickoff_async), yielding every event/result through
completely unchanged — verified against a real compiled LangGraph graph, not just a
mock. Re-attaching a fresh tracer to an already-wrapped object (a cached graph/agent
reused across requests) swaps the active tracer instead of stacking another layer of
wrapping.
One limitation: only the wrapped method is traced —
LangChainTracer traces astream_events,
not a separate ainvoke call on the same graph;
StrandsTracer traces stream_async,
not the sync agent(prompt) call.
One-shot functions
If your graph/agent/crew is rebuilt fresh on every call anyway, skip the tracer and just call one of these — same guarantees, hands back the result directly.
from agentcore_dashboard_metrics import run_agent_turn async def invoke(payload, context): # ...build graph, messages... output, turn_id = await run_agent_turn( graph, messages, log, user_id=user_id, session_id=session_id, ) return {"result": output, "turn_id": turn_id}
from agentcore_dashboard_metrics import run_strands_turn async def invoke(payload, context): agent = get_or_create_agent() output, turn_id = await run_strands_turn( agent, payload.get("prompt"), log, user_id=user_id, session_id=session_id, ) return {"result": output, "turn_id": turn_id}
from agentcore_dashboard_metrics import run_crewai_turn async def invoke(payload, context): crew = get_or_create_crew() output, turn_id = await run_crewai_turn( crew, {"prompt": payload.get("prompt")}, log, user_id=user_id, session_id=session_id, ) return {"result": output, "turn_id": turn_id}
What happens under the hood: streams the agent, marks the first
output chunk for TTFT, sums token usage across however many model calls the turn makes,
records which tool/model ran (see Call metrics below), sets status
to success or error, and emits the log line — on an exception, the line is still logged
with status="error" and the exception is re-raised unchanged.
Strands multi-agent, CrewAI Flows, or a custom loop
Drop to the framework-agnostic context manager. It handles timing and the log line; you just report usage as you get it.
from agentcore_dashboard_metrics import TurnMetrics with TurnMetrics(log, user_id=user_id, session_id=session_id) as m: result = my_own_agent_call(...) m.add_usage(input_tokens=result.input_tokens, output_tokens=result.output_tokens) m.mark_first_token() # optional — only if you stream # the log line is emitted automatically on exit, including on # an exception (status is forced to "error"; nothing is swallowed)
Already computed everything yourself? Skip straight to the formatter:
from agentcore_dashboard_metrics import log_turn_metrics log_turn_metrics( log, user_id=user_id, session_id=session_id, turn_id=turn_id, status="success", input_tokens=412, output_tokens=88, ttft_ms=640.12, latency_ms=1820.55, )
Rules that keep every panel populated
The package enforces these for you — sanitizing bad input rather than corrupting the line — but it's worth knowing what it's protecting against if you ever log this by hand.
- 01 Line starts with
Published metrics —(em dash, not a hyphen) — every panel filters on it. - 02 Fields stay in order: UserID → SessionID → TurnID → Status → InputTokens → OutputTokens → TTFTMs → LatencyMs.
- 03 Separator is exactly
" | "— space, pipe, space. - 04 Numbers are bare and unquoted —
stats sum()/avg()need real numerics. - 05
Statusis one lowercase word, no spaces or pipes. - 06 Exactly one line per turn — double-logging breaks every
count()panel. - 07 IDs are always non-empty, stable strings — they double as click-through link values.
TTFT is honest, not universal. ttft_ms
is only meaningful when the framework streams individual chunks. CrewAI's common
non-streaming kickoff_async call has no token-level stream
to time, so run_crewai_turn reports -1
there — the same convention used for any turn with no streamed output.
The Call metrics line
Published metrics gives you the turn as a whole — it can't
say which tool or model ran, or that call's own token usage. A second,
correlated line answers that:
One line per LLM step / tool call, in order (CallIndex
starting at 0), correlated back to the turn via TurnID — kept
separate rather than nested inside the turn line, since CloudWatch Logs Insights can't
unnest an array per-row, but it can group flat lines like these by name.
You get this for free with LangGraph and Strands — nothing to
change in your own code. LangGraph: run_agent_turn and
LangChainTracer record one entry per model step and per tool call.
Strands: run_strands_turn and StrandsTracer
record one entry per model cycle and per tool call, for that turn only, even when one agent
serves a whole session. CrewAI isn't wired up yet; use add_call()
below to do it yourself. Without these lines the three "Tools & Agents" panels on the turn
dashboard stay empty.
with TurnMetrics(log, user_id=user_id, session_id=session_id) as m: ... m.add_call(name="get_weather", type="tool") m.add_call(name=model_id, type="model", input_tokens=412, output_tokens=88) # one `Call metrics` line per add_call(), logged automatically on exit, # right after the turn's `Published metrics` line
add_call() is independent of add_usage()
— recording a call's own tokens does not also add them to the turn-level total; call both
if you want each.
Four dashboards, one log group
Once your agent logs this shape, these dashboards read it directly — no code changes, no per-agent wiring.
Home
Top sessions/users by token usage, system error rate, latency at a glance. Start here.
Session-based
Everything that happened in one conversation — invocation rate, TTFT, success ratio, owner.
User-based
One user's usage across all sessions — token spend, invocation trend, session list.
Turn-based
The exact detail of one exchange — latency, TTFT, tokens, status, and which tools/models ran and their own token usage.
Click-through navigation
Every ID rendered as a link jumps to the related dashboard with that value pre-filled and the time range carried over:
Getting these dashboards for your runtime
The 4 dashboards were built against one deployment, so their queries carry that deployment's AWS account id, region, runtime id and log group. One command points all 4 at yours. Everything comes from the runtime ARN you already have; no AWS credentials are read, the command only edits JSON on disk.
agentcore-dashboard-metrics create \
--arn arn:aws:bedrock-agentcore:us-east-1:111122223333:runtime/MyAgent-ab12cd34ef
# writes ./dashboards/ with all 4 JSON files, ready to import into Grafana
Or from Python: generate_dashboards(runtime_arn=...) from
agentcore_dashboard_metrics.dashboards.
Your own dashboard JSON
Export any dashboard from Grafana and point it at your runtime. There are no
placeholders to add: the package finds the account id, region, runtime id and log group
the file currently points at and swaps them for the ones in --arn.
Dashboard uids and everything else stay as they are.
agentcore-dashboard-metrics create --arn <your-runtime-arn> --template-dir my-dashboards/ --output out
- Swapped: log-group ARNs, runtime ARNs,
/aws/bedrock-agentcore/runtimes/<id>-DEFAULTnames, andaccountId/regionfields. Only string values change, never keys. - Other log groups, such as
/aws/application-signals/data, are left alone. - A file that mentions several runtimes, accounts or regions has all of them replaced. A file with nothing to replace is an error.
--log-groupsets the new log group when yours isn't AgentCore's default name.- Uids change only when you pass
--uids '{"old_uid": "new_uid"}'. For the bundled dashboards the nameshome,session,user,turnwork.
Custom log types
Register a pipe-delimited line of your own and get the matching CloudWatch Logs Insights
parse pattern from the same definition, so the log line and the
panel query can't drift apart.
from agentcore_dashboard_metrics import LogSchema, register_log_schema, log_custom_metrics schema = register_log_schema(LogSchema("Cache metrics", [ ("user_id", "UserID"), ("cache_key", "CacheKey"), ("hit", "Hit"), ])) log_custom_metrics(log, "Cache metrics", user_id="nakul", cache_key="prompt-v3", hit=True) # Cache metrics — UserID: nakul | CacheKey: prompt-v3 | Hit: True schema.parse_pattern # parse @message "Cache metrics — UserID: * | CacheKey: * | Hit: *" as user_id, cache_key, hit
API
| Name | Use for |
|---|---|
LangChainTracer(log, *, user_id, session_id).attach(graph) | LangGraph — instruments astream_events |
StrandsTracer(log, *, user_id, session_id).attach(agent) | Strands — instruments stream_async |
CrewAITracer(log, *, user_id, session_id).attach(crew) | CrewAI — instruments kickoff_async |
run_agent_turn(graph, messages, log, *, user_id, session_id, turn_id=None) | LangGraph, one-shot |
run_strands_turn(agent, prompt, log, *, user_id, session_id, turn_id=None) | Strands, one-shot |
run_crewai_turn(crew, inputs, log, *, user_id, session_id, turn_id=None) | CrewAI, one-shot |
TurnMetrics(log, *, user_id, session_id, turn_id=None) | Context manager for anything else |
log_turn_metrics(log, *, user_id, session_id, turn_id, status, input_tokens, output_tokens, ttft_ms, latency_ms) | Raw formatter — turn line |
log_call_metrics(log, *, user_id, session_id, turn_id, calls) | Raw formatter — per-call breakdown |
new_turn_id() | Generates a fresh turn id |
LogSchema(name, fields) / register_log_schema() / log_custom_metrics(log, name, **fields) | Your own log line plus its parse pattern |
dashboards.generate_dashboards(*, runtime_arn, output_dir="dashboards", ...) | Writes the 4 dashboards for your runtime |
dashboards.generate_from_templates(*, template_paths, runtime_arn, ...) | Same, for your own dashboard JSON files |
dashboards.detect_source_values(dashboard_json) | The account ids, regions, runtime ids and log groups that would be swapped |
dashboards.parse_runtime_arn(runtime_arn) | Parses an ARN into {region, account_id, runtime_id} |
All three run_*_turn() functions
are async def and return (output_text,
turn_id). On an exception, the metrics line is still logged with
status="error" and the exception re-raised — the same
guarantee applies to the tracers' wrapped methods.
"No data" checklist
- Is your agent actually logging? Confirm
Published metricslines are reaching CloudWatch at all. - Does the line match the contract in this doc? Copy one real line out of CloudWatch and run it through the dashboard's
parsepattern manually. - Is the time range wide enough? Dashboards default to
now-24h→now. - Is the variable value exact?
$SessionID/$UserID/$TurnIDdo exact string matches.