AWS Launches Framework-Agnostic Agent Evaluation via OpenTelemetry in Bedrock AgentCore
Amazon Bedrock AgentCore Evaluations now scores AI agents regardless of the framework they're built on, by reading OpenTelemetry and OpenInference instrumentation instead of requiring a specific SDK. The service automatically decodes traces from six named frameworks and extends coverage to any library following the same telemetry conventions.
What's new
Amazon has extended Bedrock AgentCore Evaluations to work with any agent framework that emits OpenTelemetry-compliant telemetry, removing a long-standing constraint where evaluation tooling only worked with agents built using a specific SDK or tracing pattern.
According to AWS, the service now supports Strands Agents, LangGraph, the OpenAI Agents SDK, LlamaIndex, Google ADK, and the Claude Agent SDK out of the box, with most frameworks covered through both OpenTelemetry and OpenInference instrumentation. Coverage also extends to any instrumentation library that follows the same naming conventions, even those not explicitly listed.
How it works
AgentCore Evaluations reconstructs an agent's session from telemetry stored in Amazon CloudWatch, routed there by the AWS Distro for OpenTelemetry (ADOT) running on Bedrock AgentCore runtime. A session is grouped by a session.id attribute; within it, each user turn is a separate trace made up of spans.
The service reads three span roles to score an interaction:
- Invoke agent spans — carry the user prompt and final agent response
- Inference spans — carry the message history and model reply for each model call
- Execute tool spans — carry the tool name, input parameters, and result
Other span types — retrieval, reranking, guardrail checks, memory operations, orchestration steps — are passed through as additional context but aren't required by the evaluators. AWS says this design is forward-compatible: unfamiliar span kinds are skipped rather than causing errors, so new span types added to future conventions won't break existing evaluation pipelines.
Framework detection happens automatically through a scope.name attribute that every OpenTelemetry instrumentation library stamps on its spans. Any library whose scope name falls under the opentelemetry.instrumentation.* or openinference.instrumentation.* prefixes is read through a generic path, without requiring configuration changes to agent code. AWS notes that a scope name outside these prefixes — even if its spans otherwise follow the conventions — will not be picked up.
Once a session is reconstructed, the same evaluators apply regardless of framework: GoalSuccessRate, Correctness, Helpfulness, and custom LLM-as-a-judge metrics score every agent identically.
Requirements
AWS lists two conditions for end-to-end evaluation. First, spans must carry a session.id attribute matching the runtimeSessionId used to invoke the agent — handled automatically by ADOT on AgentCore runtime. Second, the underlying data source must include actual message content, not just span metadata. For newly created agents using AgentCore's default unified observability setup, message content and spans live in the same log group, satisfying this requirement without extra setup.
What this means
This is a plumbing fix, not a new capability in the sense of new evaluation metrics or scoring methods. The value is in decoupling evaluation from framework lock-in: teams that mix LangGraph for orchestration, LlamaIndex for retrieval, and Claude Agent SDK for specific tasks no longer need separate evaluation pipelines or custom instrumentation for each.
The approach also signals where AWS is betting the ecosystem is headed — toward OpenTelemetry and OpenInference as de facto standards for agent observability, rather than proprietary tracing formats tied to individual SDKs. Any instrumentation author can opt into AgentCore Evaluations simply by conforming to the scope-name convention, which effectively outsources framework support to the broader open-source instrumentation community instead of requiring AWS to build and maintain per-framework adapters.
The practical limitation is that agents without OpenTelemetry-compliant instrumentation, or those using AgentCore runtime without unified observability enabled, will need configuration changes to benefit. Teams running custom or in-house agent frameworks with non-standard telemetry naming won't get automatic coverage unless they adopt the conventions.
Related Articles
Anthropic adds dynamic workflows to Claude Managed Agents, allowing up to 1,000 parallel sub-agents per execution
Anthropic has added dynamic workflows to Claude Managed Agents, letting a lead agent plan a task, distribute it to up to 1,000 parallel sub-agents per execution, and merge the results. Anthropic claims the approach found 66 of 70 hidden bugs in a 116,000-line codebase, versus 14 to 27 for a single agent. Pricing and token costs were not disclosed.
Grok Bot agent gets its own @mail.grokbot.com email address, rolling out to users now
Grok Bot, the agent available on iPhone, iPad and Mac, now has its own email address ending in @mail.grokbot.com. According to the announcement on X, the agent can use it to sign up for services, contact businesses and schedule time with people. The rollout began October 9, 2026.
Google Cloud launches Gemini agent, a multi-model enterprise agent that runs Gemini and Claude, in private preview
Google Cloud announced the Gemini agent at Gemini at Work 2026 on October 8. It is an enterprise agent that routes each job to either Gemini or Claude models. It is in private preview, with wider availability for select Workspace Business and Enterprise plans coming soon.
Microsoft's Copilot gets access to local Windows files and OS-level actions under 'Hybrid Intelligence'
Microsoft announced an upgrade to Copilot at its Windows and Surface event that gives the assistant access to local files and the ability to take actions across Windows. The company calls the underlying approach "Hybrid Intelligence," which combines local and cloud AI models. Pricing, model details, and availability were not disclosed in the available reporting.
Comments
Loading...