Natera migrated its voice‑driven phlebotomy appointment scheduler from an Amazon ECS container running a third‑party AI service to a fully managed Amazon Bedrock AgentCore runtime that employs a dual‑WebSocket bridge. The change eliminates container‑level operational work, introduces fine‑grained observability, and reduces per‑call cost while keeping latency below seven seconds.
Why the Switch to AgentCore
The original stack combined Twilio for telephony, a custom container image, and an external AI provider. Maintaining the container fleet, handling scaling events, and patching dependencies added operational burden. AgentCore abstracts the model hosting and scaling concerns, letting engineers focus on orchestration logic. Built‑in memory management lets the agent retain patient context across a call, which improves conversational flow without custom state stores. Observability is baked in: each request generates a trace that records tool invocations, inference time, and memory lookups, enabling rapid root‑cause analysis for latency or accuracy issues. The service also reports a per‑call cost under USD 0.01, which is attractive for high‑volume healthcare interactions.
Architecture: Dual‑WebSocket Bridge Pattern
The core of the new design is a dual‑WebSocket bridge. One WebSocket maintains a streaming channel with the telephony provider (Twilio in the reference implementation). A second WebSocket connects the orchestrator to the Bedrock model endpoint. The orchestrator routes audio packets between the two sockets, allowing each side to evolve independently. This separation supports swapping Twilio for Amazon Connect Health or swapping the underlying foundation model without redesigning the whole pipeline.
Two supporting techniques are highlighted:
- Event‑driven latency masking: while the model processes a turn, the orchestrator can emit interim prompts generated by a fast, lightweight model to keep the conversation responsive.
- Progressive trust model: authentication steps (personal identifiers, SMS verification) are interleaved with the dialogue, and the agent only escalates privileges after sufficient verification, reducing exposure of PHI.
Operational, Cost, and Performance Impact
Validation runs of 500 end‑to‑end calls showed 100 % tool‑calling accuracy, confirming that the agent reliably invoked downstream services (e.g., vendor availability APIs). Perceived latency stayed under seven seconds, meeting the user‑experience target for real‑time voice interactions. The cost model reported less than USD 0.01 per completed call, a notable reduction compared with the previous container‑based approach.
Observability traces from AgentCore expose three latency buckets: model inference, tool execution, and memory retrieval. Engineers can set alerts on any bucket that exceeds a threshold, enabling proactive scaling or code adjustments before patient impact.
Security and Compliance Considerations
The workflow still requires patient authentication using personal identifiers and SMS codes. The progressive trust model means that only after successful verification does the agent access or modify scheduling data, limiting the window in which PHI is exposed. Because AgentCore records detailed traces, logs must be protected in accordance with healthcare compliance regimes (e.g., HIPAA). Practitioners should ensure that trace storage is encrypted at rest and that access controls restrict log visibility to authorized personnel.
Related CloudNinjas coverage: AWS.
What This Means For Practitioners
Adopting a managed agent platform like Bedrock AgentCore can replace custom container orchestration for voice‑first use cases, delivering lower operational overhead, built‑in observability, and cost predictability. When designing similar systems, evaluate the dual‑WebSocket bridge for decoupling telephony from inference, and consider event‑driven latency masking to keep interactions snappy. Finally, map any authentication steps to a progressive trust flow and secure trace data to satisfy compliance requirements.


