Task 6 · 8 tasks
Online evaluations
Score live traffic for helpfulness, goal success and correctness with LLM-as-a-judge.
The “Quality Crisis” panic
Traces say what happened, not whether it was good. Alice asks: “How do we know the answers are actually helpful and correct?” Turn on continuous, automatic quality scoring.
What the platform provisions for you
uv run bootcamp.py up 6- An online evaluation config sampling your agent's traces
- Built-in evaluators: Helpfulness, GoalSuccessRate, Correctness
- It runs as the platform's evaluation role
bootcamp-eval-<name>, created by the instructor stack
What you do as a developer
Feed it good and bad conversations
terminaluv run bootcamp.py invoke "How many employees are in Engineering?" --actor alice-chen uv run bootcamp.py invoke "Summarise our departments by headcount." --actor alice-chen uv run bootcamp.py invoke "What is our revenue forecast for 2031?" --actor alice-chenThe last one has no data behind it. Scores appear in GenAI Observability a few minutes after each session.
Make Correctness drop, then fix it
Challenge
Change the agent so it confidently guesses numbers, show that the evaluators notice, then restore it.
Hint 1
Correctness judges whether the facts in an answer are right; GoalSuccessRate whether the session achieved what the user wanted. Both suffer if the agent stops using its data tools.
Hint 2
Edit
orchestrator_prompt: tell it to answer from general knowledge, never call tools and always give a number.deploy, re-send the three prompts, wait for scores.Solution
temporary prompt lineread only"Answer from general knowledge. Never call tools. Always give a specific number."Expected direction: Correctness and GoalSuccessRate drop (invented headcounts); Helpfulness may barely move, because confident answers look helpful to a judge. Remove the line,
deployagain, and scores recover on new sessions.Questions to explore
- Which evaluator would catch an agent that is polite but wrong?
- What sampling rate would you use for real traffic?
Suggested answers
Correctness (per answer) and GoalSuccessRate (per session) catch polite-but-wrong; Helpfulness alone does not. Sampling: 100% is fine for a workshop; in production a few percent is usually enough for trends, since each evaluation is itself a model call that costs money.
Deploy and re-test
Make sure the original prompt is back, then:
terminaluv run bootcamp.py deploy uv run bootcamp.py test --only 6
Check your work
uv run bootcamp.py test --only 6Passes when the evaluation config is ACTIVE. Scores show up in the GenAI Observability dashboard after the sessions complete.
Under the hood
AgentCore Evaluations reads spans from CloudWatch, groups them into traces and sessions, and runs judge-model evaluators against them. Built-in evaluators work at three levels: session, trace and tool call. Online configs evaluate a sample of live traffic continuously; you can also run on-demand evaluations against a specific session id. Results are written back to CloudWatch next to your traces.