Live
Self‑Managing Context in LLMs Reduces Compute Overhead and Improves ThroughputAI‑Generated OSS Vulnerability Scans Overwhelm Human Review – Implications for Security OpsBootstrapping Claude Code with Dependency Records Eliminates Initial Memory RequirementsEnterprise Copilot model control and MCP startup options in JetBrains pluginMicrosoft‑Decision‑1 Arrives on Foundry: What Engineers Need to KnowIntegrating Production Feedback into the AI Agent Lifecycle: Practical Architecture and Ops GuidanceOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model BackendSelf‑Managing Context in LLMs Reduces Compute Overhead and Improves ThroughputAI‑Generated OSS Vulnerability Scans Overwhelm Human Review – Implications for Security OpsBootstrapping Claude Code with Dependency Records Eliminates Initial Memory RequirementsEnterprise Copilot model control and MCP startup options in JetBrains pluginMicrosoft‑Decision‑1 Arrives on Foundry: What Engineers Need to KnowIntegrating Production Feedback into the AI Agent Lifecycle: Practical Architecture and Ops GuidanceOpenTelemetry tracing expands across Cloudflare’s proxy stack in betaDynamic Model Triage: Engineering Implications of Grok Bot’s Multi‑Model Backend
AWS

Replacing ECS‑Hosted Voice Scheduler with a Dual‑WebSocket Bridge on Amazon Bedrock AgentCore

AI SummaryPowered by AI

Natera moved its voice scheduling service from an ECS‑based, third‑party AI stack to a fully managed Amazon Bedrock AgentCore runtime that uses a dual‑WebSocket bridge. The shift cuts operational effort, improves latency and observability, and lowers per‑call cost, which matters to engineers building real‑time AI‑driven voice applications.

Natera migrated its voice‑driven phlebotomy appointment scheduler from an Amazon ECS container running a third‑party AI service to a fully managed Amazon Bedrock AgentCore runtime that employs a dual‑WebSocket bridge. The change eliminates container‑level operational work, introduces fine‑grained observability, and reduces per‑call cost while keeping latency below seven seconds.

Why the Switch to AgentCore

The original stack combined Twilio for telephony, a custom container image, and an external AI provider. Maintaining the container fleet, handling scaling events, and patching dependencies added operational burden. AgentCore abstracts the model hosting and scaling concerns, letting engineers focus on orchestration logic. Built‑in memory management lets the agent retain patient context across a call, which improves conversational flow without custom state stores. Observability is baked in: each request generates a trace that records tool invocations, inference time, and memory lookups, enabling rapid root‑cause analysis for latency or accuracy issues. The service also reports a per‑call cost under USD 0.01, which is attractive for high‑volume healthcare interactions.

Architecture: Dual‑WebSocket Bridge Pattern

The core of the new design is a dual‑WebSocket bridge. One WebSocket maintains a streaming channel with the telephony provider (Twilio in the reference implementation). A second WebSocket connects the orchestrator to the Bedrock model endpoint. The orchestrator routes audio packets between the two sockets, allowing each side to evolve independently. This separation supports swapping Twilio for Amazon Connect Health or swapping the underlying foundation model without redesigning the whole pipeline.

Two supporting techniques are highlighted:

  • Event‑driven latency masking: while the model processes a turn, the orchestrator can emit interim prompts generated by a fast, lightweight model to keep the conversation responsive.
  • Progressive trust model: authentication steps (personal identifiers, SMS verification) are interleaved with the dialogue, and the agent only escalates privileges after sufficient verification, reducing exposure of PHI.

Operational, Cost, and Performance Impact

Validation runs of 500 end‑to‑end calls showed 100 % tool‑calling accuracy, confirming that the agent reliably invoked downstream services (e.g., vendor availability APIs). Perceived latency stayed under seven seconds, meeting the user‑experience target for real‑time voice interactions. The cost model reported less than USD 0.01 per completed call, a notable reduction compared with the previous container‑based approach.

Observability traces from AgentCore expose three latency buckets: model inference, tool execution, and memory retrieval. Engineers can set alerts on any bucket that exceeds a threshold, enabling proactive scaling or code adjustments before patient impact.

Security and Compliance Considerations

The workflow still requires patient authentication using personal identifiers and SMS codes. The progressive trust model means that only after successful verification does the agent access or modify scheduling data, limiting the window in which PHI is exposed. Because AgentCore records detailed traces, logs must be protected in accordance with healthcare compliance regimes (e.g., HIPAA). Practitioners should ensure that trace storage is encrypted at rest and that access controls restrict log visibility to authorized personnel.

Related CloudNinjas coverage: AWS.

What This Means For Practitioners

Adopting a managed agent platform like Bedrock AgentCore can replace custom container orchestration for voice‑first use cases, delivering lower operational overhead, built‑in observability, and cost predictability. When designing similar systems, evaluate the dual‑WebSocket bridge for decoupling telephony from inference, and consider event‑driven latency masking to keep interactions snappy. Finally, map any authentication steps to a progressive trust flow and secure trace data to satisfy compliance requirements.

Originally published atAWS Machine Learning Blog