Task 5 · 8 tasks

Observability

See every model call, tool call and millisecond your agent spends, without adding code.

20 minEasy
Alice’s ask

The “Black Box” panic

“This agent serves 1,200 people and I have no idea what happens inside it!” Before something breaks, Alice wants traces: which tools ran, how long each step took, where tokens went.

Platform

What the platform provisions for you

terminal
uv run bootcamp.py up 5
  • Nothing new to create: tracing has been on since Task 4 (the agent runs under opentelemetry-instrument) and its log group already exists
  • Stage 5 switches on the check that your spans arrive in Transaction Search

Transaction Search is an account-wide switch; your instructor turned it on once for everyone.

You

What you do as a developer

  1. Generate some traffic

    terminal
    uv run bootcamp.py invoke "Which department has the most employees?" --actor alice-chen
    uv run bootcamp.py invoke "Headcount in Sales and the weather in Chicago?" --actor alice-chen
  2. Find where the time goes

    Challenge

    For the second question, find the slowest span and decide what dominates latency: model calls, the Gateway tool call, or the weather API.

    Hint 1

    In the AWS console open CloudWatch → GenAI Observability → Bedrock AgentCore, pick your agent, then a session and its trace. The trace view shows spans as a timeline.

    Hint 2

    Expand the data_agent and weather_agent tool spans: each contains the specialist's own model calls. Compare those with the Gateway query_db span inside.

    Solution

    Usually the model calls dominate: the orchestrator's routing call plus each specialist's calls, several seconds in total. The Gateway query_db span is typically well under a second, and the weather lookup is a couple of HTTP requests. Specialists run one after another, so their time adds up. Levers: fewer model round-trips (better tool docstrings), a faster model, or running specialists in parallel.

  3. Experiments

    Switch MODEL_ID in .env, deploy, send the same prompts and compare latency and token counts in the traces.

    What should I expect?

    claude-sonnet-5-5 tends to route with fewer, better tool calls but each call is slower and pricier; gpt-6-luna is faster per call but may need extra round-trips. Measure, don't guess.

Check your work

terminal
uv run bootcamp.py test --only 5

Spans can take about 10 minutes to appear. A WARN here just means “not yet”; re-run later.

Under the hood

AgentCore Runtime ships the AWS Distro for OpenTelemetry. Strands emits OpenTelemetry spans for agent cycles, model calls and tool calls, and the runtime exports them to X-Ray. With Transaction Search enabled, spans land in the aws/spans log group, where the GenAI Observability dashboards and the evaluations in Task 6 read them. Application logs go to /aws/bedrock-agentcore/runtimes/<runtime-id>-DEFAULT.