Live
Self‑Managing Context in LLMs Reduces Compute Overhead and Improves ThroughputAI‑Generated OSS Vulnerability Scans Overwhelm Human Review – Implications for Security OpsBootstrapping Claude Code with Dependency Records Eliminates Initial Memory RequirementsEnterprise Copilot model control and MCP startup options in JetBrains pluginMicrosoft‑Decision‑1 Arrives on Foundry: What Engineers Need to KnowIntegrating Production Feedback into the AI Agent Lifecycle: Practical Architecture and Ops GuidanceOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model BackendSelf‑Managing Context in LLMs Reduces Compute Overhead and Improves ThroughputAI‑Generated OSS Vulnerability Scans Overwhelm Human Review – Implications for Security OpsBootstrapping Claude Code with Dependency Records Eliminates Initial Memory RequirementsEnterprise Copilot model control and MCP startup options in JetBrains pluginMicrosoft‑Decision‑1 Arrives on Foundry: What Engineers Need to KnowIntegrating Production Feedback into the AI Agent Lifecycle: Practical Architecture and Ops GuidanceOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model Backend
AWS

OpenTelemetry‑Based Evaluation Unifies Bedrock AgentCore Traces Across Frameworks

AI SummaryPowered by AI

Amazon Bedrock AgentCore now evaluates agents by consuming OpenTelemetry spans, removing the need for framework‑specific evaluation adapters. This change lets engineers keep their preferred agent libraries while ensuring consistent scoring and observability.

Amazon Bedrock AgentCore’s evaluation service now bases its scoring on OpenTelemetry traces instead of a particular SDK or client library. By reading a standard set of span types, the service can evaluate agents built with LangGraph, LlamaIndex, OpenAI Agents SDK, Google ADK, Claude Agent SDK, Strands Agents, or any future framework that emits OpenTelemetry data.

How the Evaluation Service Consumes Traces

The service expects three core span roles:

  • Invoke‑agent span – the top‑level request/response for a single user turn; it carries the raw user prompt and the final agent reply.
  • Inference span – each LLM call; it includes the message history sent to the model and the model’s output.
  • Execute‑tool span – a tool invocation; it records the tool name, input parameters, and result.

All other spans – retrieval, reranking, guardrails, memory, orchestration – are ignored by the evaluator but remain visible in the trace for context. The service classifies incoming spans using the OpenTelemetry GenAI semantic conventions (chat, embeddings, retrieval, etc.) and the OpenInference schema (LLM, TOOL, RETRIEVER, …). Unrecognised span kinds are simply skipped, preserving forward compatibility.

Session (runtimeSessionId)
└── Trace (user turn, trace_id)
    ├── invoke agent span ← user prompt + final response
    ├── inference span ← messages + model reply
    ├── execute tool span ← tool name + params + result
    ├── retriever span (context only)
    ├── inference span ← next model call with tool result
    └── … (guardrail, memory, etc.)

Impact on Architecture and Implementation

Teams no longer need a dedicated evaluation adapter for each framework. The only architectural requirement is that the agent’s runtime emits OpenTelemetry spans over the OpenTelemetry Protocol (OTLP). On AgentCore, the AWS Distro for OpenTelemetry (ADOT) collects those spans and forwards them to Amazon CloudWatch, where the evaluation service reads them.

Practically, this means:

  • Existing codebases can keep their chosen orchestration or retrieval library unchanged.
  • Instrumentation can be added via the native OpenTelemetry SDK or a community library that maps framework‑specific events to the GenAI/ OpenInference attributes.
  • Future framework updates that introduce new span kinds will not break the evaluation pipeline, because the service ignores unknown kinds.

Operational Considerations

Because evaluation now lives on the same telemetry pipeline that feeds CloudWatch, operators should monitor the health of the OTLP exporter and ADOT agents. Latency in span delivery could delay evaluation results, and missing spans (e.g., a missing invoke‑agent span) will cause the evaluator to skip scoring that session.

Scaling the evaluation service is independent of the agent runtime; the service processes spans as they arrive in CloudWatch. Teams can therefore size their AgentCore instances based on request throughput without allocating extra resources for a bespoke evaluation harness.

Security and Data‑Handling Implications

Telemetry streams contain raw user prompts, model outputs, and tool results. Once exported, these records reside in CloudWatch logs and metrics. Practitioners should treat the telemetry store as a repository of potentially sensitive data and apply appropriate IAM policies, log retention settings, and audit controls.

Because the evaluation service only reads the three required span roles, any additional spans that carry proprietary data are not used for scoring, but they are still persisted. Organizations may need to review what extra context is being emitted and whether it aligns with data‑privacy requirements.

Related CloudNinjas coverage: AWS.

What This Means For Practitioners

  • Verify that your agent framework is instrumented with OpenTelemetry and that spans are exported via OTLP to ADOT.
  • Confirm that the three required span roles (invoke‑agent, inference, execute‑tool) include the documented attributes; otherwise evaluation will be incomplete.
  • Review CloudWatch access controls and retention policies to protect prompt and response data.
  • Monitor the OTLP exporter health to avoid gaps in evaluation coverage.
  • Plan for future framework changes knowing that unknown span kinds will be ignored, not cause failures.
Originally published atAWS Machine Learning Blog