Amazon Bedrock AgentCore’s evaluation service now bases its scoring on OpenTelemetry traces instead of a particular SDK or client library. By reading a standard set of span types, the service can evaluate agents built with LangGraph, LlamaIndex, OpenAI Agents SDK, Google ADK, Claude Agent SDK, Strands Agents, or any future framework that emits OpenTelemetry data.
How the Evaluation Service Consumes Traces
The service expects three core span roles:
- Invoke‑agent span – the top‑level request/response for a single user turn; it carries the raw user prompt and the final agent reply.
- Inference span – each LLM call; it includes the message history sent to the model and the model’s output.
- Execute‑tool span – a tool invocation; it records the tool name, input parameters, and result.
All other spans – retrieval, reranking, guardrails, memory, orchestration – are ignored by the evaluator but remain visible in the trace for context. The service classifies incoming spans using the OpenTelemetry GenAI semantic conventions (chat, embeddings, retrieval, etc.) and the OpenInference schema (LLM, TOOL, RETRIEVER, …). Unrecognised span kinds are simply skipped, preserving forward compatibility.
Session (runtimeSessionId)
└── Trace (user turn, trace_id)
├── invoke agent span ← user prompt + final response
├── inference span ← messages + model reply
├── execute tool span ← tool name + params + result
├── retriever span (context only)
├── inference span ← next model call with tool result
└── … (guardrail, memory, etc.)
Impact on Architecture and Implementation
Teams no longer need a dedicated evaluation adapter for each framework. The only architectural requirement is that the agent’s runtime emits OpenTelemetry spans over the OpenTelemetry Protocol (OTLP). On AgentCore, the AWS Distro for OpenTelemetry (ADOT) collects those spans and forwards them to Amazon CloudWatch, where the evaluation service reads them.
Practically, this means:
- Existing codebases can keep their chosen orchestration or retrieval library unchanged.
- Instrumentation can be added via the native OpenTelemetry SDK or a community library that maps framework‑specific events to the GenAI/ OpenInference attributes.
- Future framework updates that introduce new span kinds will not break the evaluation pipeline, because the service ignores unknown kinds.
Operational Considerations
Because evaluation now lives on the same telemetry pipeline that feeds CloudWatch, operators should monitor the health of the OTLP exporter and ADOT agents. Latency in span delivery could delay evaluation results, and missing spans (e.g., a missing invoke‑agent span) will cause the evaluator to skip scoring that session.
Scaling the evaluation service is independent of the agent runtime; the service processes spans as they arrive in CloudWatch. Teams can therefore size their AgentCore instances based on request throughput without allocating extra resources for a bespoke evaluation harness.
Security and Data‑Handling Implications
Telemetry streams contain raw user prompts, model outputs, and tool results. Once exported, these records reside in CloudWatch logs and metrics. Practitioners should treat the telemetry store as a repository of potentially sensitive data and apply appropriate IAM policies, log retention settings, and audit controls.
Because the evaluation service only reads the three required span roles, any additional spans that carry proprietary data are not used for scoring, but they are still persisted. Organizations may need to review what extra context is being emitted and whether it aligns with data‑privacy requirements.
Related CloudNinjas coverage: AWS.
What This Means For Practitioners
- Verify that your agent framework is instrumented with OpenTelemetry and that spans are exported via OTLP to ADOT.
- Confirm that the three required span roles (invoke‑agent, inference, execute‑tool) include the documented attributes; otherwise evaluation will be incomplete.
- Review CloudWatch access controls and retention policies to protect prompt and response data.
- Monitor the OTLP exporter health to avoid gaps in evaluation coverage.
- Plan for future framework changes knowing that unknown span kinds will be ignored, not cause failures.


