Task 6 · 8 tasks

Online evaluations

Score live traffic for helpfulness, goal success and correctness with LLM-as-a-judge.

20 minMedium
Alice’s ask

The “Quality Crisis” panic

Traces say what happened, not whether it was good. Alice asks: “How do we know the answers are actually helpful and correct?” Turn on continuous, automatic quality scoring.

Platform

What the platform provisions for you

terminal
uv run bootcamp.py up 6
  • An online evaluation config sampling your agent's traces
  • Built-in evaluators: Helpfulness, GoalSuccessRate, Correctness
  • It runs as the platform's evaluation role bootcamp-eval-<name>, created by the instructor stack
You

What you do as a developer

  1. Feed it good and bad conversations

    terminal
    uv run bootcamp.py invoke "How many employees are in Engineering?" --actor alice-chen
    uv run bootcamp.py invoke "Summarise our departments by headcount." --actor alice-chen
    uv run bootcamp.py invoke "What is our revenue forecast for 2031?" --actor alice-chen

    The last one has no data behind it. Scores appear in GenAI Observability a few minutes after each session.

  2. Make Correctness drop, then fix it

    Challenge

    Change the agent so it confidently guesses numbers, show that the evaluators notice, then restore it.

    Hint 1

    Correctness judges whether the facts in an answer are right; GoalSuccessRate whether the session achieved what the user wanted. Both suffer if the agent stops using its data tools.

    Hint 2

    Edit orchestrator_prompt: tell it to answer from general knowledge, never call tools and always give a number. deploy, re-send the three prompts, wait for scores.

    Solution
    temporary prompt lineread only
    "Answer from general knowledge. Never call tools. Always give a specific number."

    Expected direction: Correctness and GoalSuccessRate drop (invented headcounts); Helpfulness may barely move, because confident answers look helpful to a judge. Remove the line, deploy again, and scores recover on new sessions.

  3. Questions to explore

    • Which evaluator would catch an agent that is polite but wrong?
    • What sampling rate would you use for real traffic?
    Suggested answers

    Correctness (per answer) and GoalSuccessRate (per session) catch polite-but-wrong; Helpfulness alone does not. Sampling: 100% is fine for a workshop; in production a few percent is usually enough for trends, since each evaluation is itself a model call that costs money.

  4. Deploy and re-test

    Make sure the original prompt is back, then:

    terminal
    uv run bootcamp.py deploy
    uv run bootcamp.py test --only 6

Check your work

terminal
uv run bootcamp.py test --only 6

Passes when the evaluation config is ACTIVE. Scores show up in the GenAI Observability dashboard after the sessions complete.

Under the hood

AgentCore Evaluations reads spans from CloudWatch, groups them into traces and sessions, and runs judge-model evaluators against them. Built-in evaluators work at three levels: session, trace and tool call. Online configs evaluate a sample of live traffic continuously; you can also run on-demand evaluations against a specific session id. Results are written back to CloudWatch next to your traces.