View a markdown version of this page

实地真相评估 - 亚马逊基岩 AgentCore

本文属于机器翻译版本。若本译文内容与英语原文存在差异,则一律以英文原文为准。

实地真相评估

基本事实是给定输入的已知正确答案或预期行为——这是你将实际结果与之比较的 “黄金标准”。对于代理评估,ground truth 将主观质量评估转化为客观衡量标准,从而实现回归检测、基准数据集和特定领域的正确性,这是通用评估者无法提供的。

通过实况评估,您可以在调用评估 API 时在会话跨度旁边提供参考输入。该服务使用这些参考输入对您的代理的实际行为和预期行为进行评分。不使用特定实况字段的评估人员会忽略该字段,并报告响应中未使用哪些字段。

支持的内置评估器和实况字段

下表显示了哪些内置评估器支持实情以及它们使用的字段。

评估员 级别 实地真相字段 说明

Builtin.Correctness

追踪

expectedResponse

衡量代理的响应与预期答案匹配的准确程度。使用 LLM-as-a-Judge 得分。

Builtin.GoalSuccessRate

会话

assertions

验证代理的行为是否满足整个会话中的自然语言断言。使用 LLM-as-a-Judge 得分。

Builtin.TrajectoryExactOrderMatch

会话

expectedTrajectory

检查实际的工具调用顺序是否与预期的顺序完全匹配——相同的工具,相同的顺序,没有额外内容。程序化评分(无 LLM 调用)。

Builtin.TrajectoryInOrderMatch

会话

expectedTrajectory

检查所有预期工具是否在实际序列中按顺序显示,但允许在它们之间使用额外的工具。程序化评分。

Builtin.TrajectoryAnyOrderMatch

会话

expectedTrajectory

检查所有预期的工具是否都存在于实际序列中,无论顺序如何。允许使用额外的工具。程序化评分。

注意

自定义评估器还通过评估说明中的占位符支持基本真值字段。有关详细信息,请参阅自定义评估器中的基本真相。

下表描述了基本真值字段。

字段 Type Scope 说明

expectedResponse

字符串

追踪

特定回合的预期代理响应。在参考输入上下文traceId中使用的范围限定为跟踪。

assertions

字符串列表

会话

关于代理在整个会话中的行为的自然语言陈述,应该是真实的。

expectedTrajectory

工具名称清单

会话

会话的预期工具调用顺序。

  • 实况字段是可选的。如果你省略它们,评估器就会退回到其基本的无真值模式(例如,在没有真相的情况下Builtin.Correctness仍然可以工作expectedResponse,它仅根据上下文进行评估)。

  • 您可以在单个请求中提供所有实情字段。该服务会为每个评估者选择相关字段,并在响应ignoredReferenceInputFields中报告任何未使用的字段。

  • 您无需提供expectedResponse每条线索。使用评估器的无地面真值变体来评估没有地面真值的痕迹。

先决条件

  • Python 3.10+

  • 使用支持的框架和仪器库构建的代理。有关支持的框架和工具库的更多信息,请参阅支持的代理框架。

  • 在 AgentCore Runtime 上部署且启用了可观察性的代理,或者使用配置了可AgentCore 观测性(包括事务搜索)的受支持框架构建的代理。有关遥测设置的更多信息,请参阅遥测设置和交付。

  • AWS 使用bedrock-agentcorebedrock-agentcore-control、和 logs (CloudWatch) 的权限配置的证书

有关下载会话跨度的说明,请参阅按需评估按需评估入门入门。

关于示例

本页上的示例使用AgentCore 评估教程中的示例代理。该代理有两个工具 weather —— calculator 和,部署在 AgentCore Runtime 上,启用了可观察性。

这些示例假设会话为两回合:

  1. 回合 1:“什么是 15 + 27?” — 代理使用该calculator工具并用结果进行响应。

  2. 回合 2:“天气怎么样?” — 代理使用该weather工具并根据当前天气进行响应。

在运行评估之前,请调用您的代理并等待 2-5 分钟 CloudWatch 以提取遥测数据。

本页的示例中使用了以下常量。用你自己的值替换它们:

REGION = "<region-code>" AGENT_ID = "my-agent-id" SESSION_ID = "my-session-id" TRACE_ID_1 = "<trace-id-1>" # Turn 1: "What is 15 + 27?" TRACE_ID_2 = "<trace-id-2>" # Turn 2: "What's the weather?"

正确性与预期的响应

Builtin.Correctness是一个跟踪级别评估器,用于衡量代理的响应与预期答案匹配的准确程度。当你提供时expectedResponse,评估人员会使用 LLM-as-a-Judge 评分将代理的实际反应与你的真实情况进行比较。

例
AgentCore SDK
  1. from bedrock_agentcore.evaluation import EvaluationClient, ReferenceInputs client = EvaluationClient(region_name=REGION) # String form — matched against the last trace in the session results = client.run( evaluator_ids=["Builtin.Correctness"], agent_id=AGENT_ID, session_id=SESSION_ID, reference_inputs=ReferenceInputs( expected_response="The weather is sunny", ), ) for r in results: print(f"Trace: {r['context']['spanContext'].get('traceId', 'session')}") print(f"Score: {r['value']}, Label: {r['label']}")

    要定位特定的跟踪,请将跟踪 ID 映射到预期答案的字典expected_response形式传递:

    results = client.run( evaluator_ids=["Builtin.Correctness"], agent_id=AGENT_ID, session_id=SESSION_ID, reference_inputs=ReferenceInputs( expected_response={ TRACE_ID_1: "15 + 27 = 42", TRACE_ID_2: "The weather is sunny", }, ), )
AgentCore CLI
  1. # Expected response matched against the last trace agentcore run eval \ --runtime AGENT_NAME \ --session-id SESSION_ID \ --evaluator Builtin.Correctness \ --expected-response "The weather is sunny" # Target a specific trace agentcore run eval \ --runtime AGENT_NAME \ --session-id SESSION_ID \ --evaluator Builtin.Correctness \ --trace-id TRACE_ID_1 \ --expected-response "15 + 27 = 42" # ARN mode — evaluate an agent outside the CLI project agentcore run eval \ --runtime-arn arn:aws:bedrock-agentcore:<region-code>:<account-id>:runtime/<agent-id> \ --region <region-code> \ --session-id SESSION_ID \ --evaluator-arn arn:aws:bedrock-agentcore:::evaluator/Builtin.Correctness \ --expected-response "The weather is sunny"
Starter Toolkit SDK
  1. from bedrock_agentcore_starter_toolkit import Evaluation, ReferenceInputs eval_client = Evaluation(region=REGION) # String form — matched against the last trace results = eval_client.run( agent_id=AGENT_ID, session_id=SESSION_ID, evaluators=["Builtin.Correctness"], reference_inputs=ReferenceInputs( expected_response="The weather is sunny", ), ) for r in results.get_successful_results(): print(f"Score: {r.value:.2f}, Label: {r.label}")

    要定位特定的跟踪,请传递一个元组:(trace_id, expected_response)

    results = eval_client.run( agent_id=AGENT_ID, session_id=SESSION_ID, evaluators=["Builtin.Correctness"], reference_inputs=ReferenceInputs( expected_response=(TRACE_ID_1, "15 + 27 = 42"), ), )
AgentCore CLI
  1. # Expected response matched against the last trace agentcore run eval \ --runtime-arn AGENT_RUNTIME_ARN \ --region REGION \ --session-id SESSION_ID \ --evaluator-arn arn:aws:bedrock-agentcore:::evaluator/Builtin.Correctness \ --expected-response "The weather is sunny" # Target a specific trace agentcore run eval \ --runtime-arn AGENT_RUNTIME_ARN \ --region REGION \ --session-id SESSION_ID \ --trace-id TRACE_ID_1 \ --evaluator-arn arn:aws:bedrock-agentcore:::evaluator/Builtin.Correctness \ --expected-response "15 + 27 = 42" # Save results to a file agentcore run eval \ --runtime-arn AGENT_RUNTIME_ARN \ --region REGION \ --session-id SESSION_ID \ --evaluator-arn arn:aws:bedrock-agentcore:::evaluator/Builtin.Correctness \ --expected-response "The weather is sunny" \ --output results.json
AWS SDK (boto3)
  1. import boto3 client = boto3.client("bedrock-agentcore", region_name=REGION) response = client.evaluate( evaluatorId="Builtin.Correctness", evaluationInput={"sessionSpans": session_spans_and_log_events}, evaluationReferenceInputs=[ { "context": { "spanContext": { "sessionId": SESSION_ID, "traceId": TRACE_ID_1 } }, "expectedResponse": {"text": "15 + 27 = 42"} }, { "context": { "spanContext": { "sessionId": SESSION_ID, "traceId": TRACE_ID_2 } }, "expectedResponse": {"text": "The weather is sunny"} } ] ) for result in response["evaluationResults"]: print(f"Score: {result['value']}, Label: {result['label']}")

GoalSuccessRate 附带断言

Builtin.GoalSuccessRate是一个会话级评估器,用于验证代理的行为是否满足一组自然语言断言。断言可以检查整个对话中的工具使用情况、响应内容、操作顺序或任何其他可观察到的行为。

注意

以下示例使用断言来验证工具的使用情况,但断言是自由形式的自然语言——您可以使用它们来断言代理行为的任何方面,例如回应语气、事实准确性、安全合规性或业务逻辑。

例
AgentCore SDK
  1. from bedrock_agentcore.evaluation import EvaluationClient, ReferenceInputs client = EvaluationClient(region_name=REGION) results = client.run( evaluator_ids=["Builtin.GoalSuccessRate"], agent_id=AGENT_ID, session_id=SESSION_ID, reference_inputs=ReferenceInputs( assertions=[ "Agent used the calculator tool to compute the result", "Agent returned the correct numerical answer of 42", "Agent used the weather tool when asked about weather", ], ), ) for r in results: print(f"Score: {r['value']}, Label: {r['label']}") print(f"Explanation: {r['explanation'][:200]}")
AgentCore CLI
  1. agentcore run eval \ --runtime AGENT_NAME \ --session-id SESSION_ID \ --evaluator Builtin.GoalSuccessRate \ --assertion "Agent used the calculator tool to compute the result" \ --assertion "Agent returned the correct numerical answer of 42" \ --assertion "Agent used the weather tool when asked about weather" # ARN mode — evaluate an agent outside the CLI project agentcore run eval \ --runtime-arn arn:aws:bedrock-agentcore:<region-code>:<account-id>:runtime/<agent-id> \ --region <region-code> \ --session-id SESSION_ID \ --evaluator-arn arn:aws:bedrock-agentcore:::evaluator/Builtin.GoalSuccessRate \ --assertion "Agent used the calculator tool to compute the result" \ --assertion "Agent returned the correct numerical answer of 42"
Starter Toolkit SDK
  1. from bedrock_agentcore_starter_toolkit import Evaluation, ReferenceInputs eval_client = Evaluation(region=REGION) results = eval_client.run( agent_id=AGENT_ID, session_id=SESSION_ID, evaluators=["Builtin.GoalSuccessRate"], reference_inputs=ReferenceInputs( assertions=[ "Agent used the calculator tool to compute the result", "Agent returned the correct numerical answer of 42", "Agent used the weather tool when asked about weather", ], ), ) for r in results.get_successful_results(): print(f"Score: {r.value:.2f}, Label: {r.label}")
AgentCore CLI
  1. agentcore run eval \ --runtime-arn AGENT_RUNTIME_ARN \ --region REGION \ --session-id SESSION_ID \ --evaluator-arn arn:aws:bedrock-agentcore:::evaluator/Builtin.GoalSuccessRate \ --assertion "Agent used the calculator tool to compute the result" \ --assertion "Agent returned the correct numerical answer of 42" \ --assertion "Agent used the weather tool when asked about weather"
AWS SDK (boto3)
  1. import boto3 client = boto3.client("bedrock-agentcore", region_name=REGION) response = client.evaluate( evaluatorId="Builtin.GoalSuccessRate", evaluationInput={"sessionSpans": session_spans_and_log_events}, evaluationReferenceInputs=[ { "context": { "spanContext": { "sessionId": SESSION_ID } }, "assertions": [ {"text": "Agent used the calculator tool to compute the result"}, {"text": "Agent returned the correct numerical answer of 42"}, {"text": "Agent used the weather tool when asked about weather"} ] } ] ) for result in response["evaluationResults"]: print(f"Score: {result['value']}, Label: {result['label']}")

轨迹与预期轨迹匹配

轨迹评估器将代理的实际工具调用顺序与预期的工具名称序列进行比较。有三种变体可供选择,每种变体都有不同的匹配严格度。这三个都是会话级评估器,使用编程评分(没有 LLM 调用,因此代币使用量为零)。

评估员 匹配规则 示例

Builtin.TrajectoryExactOrderMatch

实际必须完全符合预期 — 相同的工具,相同的顺序,没有额外费用

预期:[calculator, weather],实际:[calculator, weather]→ 通过。实际:[calculator, weather, calculator]→ 失败。

Builtin.TrajectoryInOrderMatch

预期的工具必须按顺序出现,但允许在它们之间使用额外的工具

预期:[calculator, weather],实际:[calculator, some_tool, weather]→ 通过。

Builtin.TrajectoryAnyOrderMatch

所有预期的工具都必须在场,顺序无关紧要,允许额外使用

预期:[calculator, weather],实际:[weather, calculator]→ 通过。

例
AgentCore SDK
  1. from bedrock_agentcore.evaluation import EvaluationClient, ReferenceInputs client = EvaluationClient(region_name=REGION) results = client.run( evaluator_ids=[ "Builtin.TrajectoryExactOrderMatch", "Builtin.TrajectoryInOrderMatch", "Builtin.TrajectoryAnyOrderMatch", ], agent_id=AGENT_ID, session_id=SESSION_ID, reference_inputs=ReferenceInputs( expected_trajectory=["calculator", "weather"], ), ) for r in results: print(f"{r['evaluatorId']}: {r['value']} ({r['label']})") print(f" {r['explanation'][:150]}")
AgentCore CLI
  1. 工具名称以逗号分隔的列表形式传递:

    agentcore run eval \ --runtime AGENT_NAME \ --session-id SESSION_ID \ --evaluator Builtin.TrajectoryExactOrderMatch Builtin.TrajectoryInOrderMatch Builtin.TrajectoryAnyOrderMatch \ --expected-trajectory "calculator,weather" # ARN mode — evaluate an agent outside the CLI project agentcore run eval \ --runtime-arn arn:aws:bedrock-agentcore:<region-code>:<account-id>:runtime/<agent-id> \ --region <region-code> \ --session-id SESSION_ID \ --evaluator-arn arn:aws:bedrock-agentcore:::evaluator/Builtin.TrajectoryExactOrderMatch \ --expected-trajectory "calculator,weather"
Starter Toolkit SDK
  1. from bedrock_agentcore_starter_toolkit import Evaluation, ReferenceInputs eval_client = Evaluation(region=REGION) results = eval_client.run( agent_id=AGENT_ID, session_id=SESSION_ID, evaluators=[ "Builtin.TrajectoryExactOrderMatch", "Builtin.TrajectoryInOrderMatch", "Builtin.TrajectoryAnyOrderMatch", ], reference_inputs=ReferenceInputs( expected_trajectory=["calculator", "weather"], ), ) for r in results.get_successful_results(): print(f"{r.evaluator_name}: {r.value:.2f} ({r.label})")
AgentCore CLI
  1. 工具名称以逗号分隔的列表形式传递:

    agentcore run eval \ --runtime-arn AGENT_RUNTIME_ARN \ --region REGION \ --session-id SESSION_ID \ --evaluator-arn arn:aws:bedrock-agentcore:::evaluator/Builtin.TrajectoryExactOrderMatch arn:aws:bedrock-agentcore:::evaluator/Builtin.TrajectoryInOrderMatch arn:aws:bedrock-agentcore:::evaluator/Builtin.TrajectoryAnyOrderMatch \ --expected-trajectory "calculator,weather"
AWS SDK (boto3)
  1. import boto3 client = boto3.client("bedrock-agentcore", region_name=REGION) for evaluator in [ "Builtin.TrajectoryExactOrderMatch", "Builtin.TrajectoryInOrderMatch", "Builtin.TrajectoryAnyOrderMatch", ]: response = client.evaluate( evaluatorId=evaluator, evaluationInput={"sessionSpans": session_spans_and_log_events}, evaluationReferenceInputs=[ { "context": { "spanContext": { "sessionId": SESSION_ID } }, "expectedTrajectory": { "toolNames": ["calculator", "weather"] } } ] ) for result in response["evaluationResults"]: print(f"{result['evaluatorId']}: {result['value']} ({result['label']})")

将所有实况字段合并到一个请求中

您可以在一次评估调用中将所有实况字段一起传递。该服务将每个字段路由到相应的赋值器,并忽略给定评估者未使用的字段。这意味着您只需构造一次参考输入,即可在不同的评估器中重复使用它们,而无需修改负载。

例
AgentCore SDK
  1. from bedrock_agentcore.evaluation import EvaluationClient, ReferenceInputs client = EvaluationClient(region_name=REGION) results = client.run( evaluator_ids=[ "Builtin.Correctness", "Builtin.GoalSuccessRate", "Builtin.TrajectoryExactOrderMatch", "Builtin.TrajectoryInOrderMatch", "Builtin.TrajectoryAnyOrderMatch", ], agent_id=AGENT_ID, session_id=SESSION_ID, reference_inputs=ReferenceInputs( expected_response="The weather is sunny", assertions=[ "Agent used the calculator tool for math", "Agent used the weather tool when asked about weather", ], expected_trajectory=["calculator", "weather"], ), ) for r in results: ignored = r.get("ignoredReferenceInputFields", []) print(f"{r['evaluatorId']}: {r['value']} ({r['label']})") if ignored: print(f" Ignored fields: {ignored}")
AgentCore CLI
  1. agentcore run eval \ --runtime AGENT_NAME \ --session-id SESSION_ID \ --evaluator Builtin.Correctness Builtin.GoalSuccessRate Builtin.TrajectoryExactOrderMatch \ --assertion "Agent used the calculator tool for math" \ --assertion "Agent used the weather tool when asked about weather" \ --expected-trajectory "calculator,weather" \ --expected-response "The weather is sunny" \ --output results.json
Starter Toolkit SDK
  1. from bedrock_agentcore_starter_toolkit import Evaluation, ReferenceInputs eval_client = Evaluation(region=REGION) results = eval_client.run( agent_id=AGENT_ID, session_id=SESSION_ID, evaluators=[ "Builtin.Correctness", "Builtin.GoalSuccessRate", "Builtin.TrajectoryExactOrderMatch", "Builtin.TrajectoryInOrderMatch", "Builtin.TrajectoryAnyOrderMatch", ], reference_inputs=ReferenceInputs( expected_response="The weather is sunny", assertions=[ "Agent used the calculator tool for math", "Agent used the weather tool when asked about weather", ], expected_trajectory=["calculator", "weather"], ), ) for r in results.get_successful_results(): print(f"{r.evaluator_name}: {r.value:.2f} ({r.label})")
AWS SDK (boto3)
  1. import boto3 client = boto3.client("bedrock-agentcore", region_name=REGION) reference_inputs = [ { "context": { "spanContext": {"sessionId": SESSION_ID} }, "assertions": [ {"text": "Agent used the calculator tool for math"}, {"text": "Agent used the weather tool when asked about weather"} ], "expectedTrajectory": { "toolNames": ["calculator", "weather"] } }, { "context": { "spanContext": { "sessionId": SESSION_ID, "traceId": TRACE_ID_2 } }, "expectedResponse": {"text": "The weather is sunny"} } ] for evaluator in ["Builtin.Correctness", "Builtin.GoalSuccessRate", "Builtin.TrajectoryExactOrderMatch"]: response = client.evaluate( evaluatorId=evaluator, evaluationInput={"sessionSpans": session_spans_and_log_events}, evaluationReferenceInputs=reference_inputs ) for result in response["evaluationResults"]: ignored = result.get("ignoredReferenceInputFields", []) print(f"{result['evaluatorId']}: {result['value']} ({result['label']})") if ignored: print(f" Ignored fields: {ignored}")

了解忽略的参考输入字段

当您提供评估者不使用的事实字段时,响应中会包含一个ignoredReferenceInputFields列出未使用字段的数组。这是信息性的,不是错误的——评估仍然成功完成。

例如,如果您使用 provided Builtin.Helpfulness 进行expectedResponse调用,则评估者会忽略基本事实(Helpfulness 不使用它)并返回:

{ "evaluatorId": "Builtin.Helpfulness", "value": 0.83, "label": "Very Helpful", "explanation": "...", "ignoredReferenceInputFields": ["expectedResponse"] }

这种行为是设计使然,它允许您构造一组参考输入,并在多个评估器中使用它们,而无需调整每个评估器的有效载荷。

自定义评估器中的真实情况

自定义评估者可以在评估说明中通过占位符使用实况字段。创建自定义评估器时,可以引用以下占位符:

  • Session-level 自定义评估器:{context}、、{available_tools}、{actual_tool_trajectory}、{expected_tool_trajectory} {assertions}

  • Trace-level 自定义评估器:{context},, {assistant_turn} {expected_response}

例如,检查响应相似度的自定义跟踪级别评估器可能会使用:

Compare the agent's response with the expected response. Agent response: {assistant_turn} Expected response: {expected_response} Rate how closely the agent's response matches the expected response on a scale of 0 to 1.

当在参考输入expectedResponse中调用此评估器时,服务会在评分之前用实际的真实值替换占位符。

有关创建自定义评估器的详细信息,请参阅自定义评估器。

注意

使用实况占位符 ({assertions},{expected_response},{expected_tool_trajectory}) 的自定义评估器不能用于在线评估配置,因为在线评估会监控没有地面真实值的实时制作流量。