Live
Aurora PostgreSQL adds native Iceberg and Parquet querying via DuckDBAI‑Driven Vulnerability Discovery: Rising Volume and Faster Exploitation Demand New Ops PracticesCISO Alignment for Cybersecurity Startups: Engineering Practices That Win Security LeadershipHydraFusion multi‑model orchestration lands in VS Code and Copilot appUsing Bedrock Knowledge Bases for RAG‑Based Claim LookupDeploying Multi‑Agent Workflows on Amazon Bedrock AgentCore Runtime InstancesAI vulnerability benchmark from AWS reveals stubborn false‑positive ratesCoreWeave Deploys NVIDIA Vera Rubin GPUs and Vera CPUs for Scalable Agentic AI WorkloadsAurora PostgreSQL adds native Iceberg and Parquet querying via DuckDBAI‑Driven Vulnerability Discovery: Rising Volume and Faster Exploitation Demand New Ops PracticesCISO Alignment for Cybersecurity Startups: Engineering Practices That Win Security LeadershipHydraFusion multi‑model orchestration lands in VS Code and Copilot appUsing Bedrock Knowledge Bases for RAG‑Based Claim LookupDeploying Multi‑Agent Workflows on Amazon Bedrock AgentCore Runtime InstancesAI vulnerability benchmark from AWS reveals stubborn false‑positive ratesCoreWeave Deploys NVIDIA Vera Rubin GPUs and Vera CPUs for Scalable Agentic AI Workloads
AWS

Deploying Multi‑Agent Workflows on Amazon Bedrock AgentCore Runtime Instances

AI SummaryPowered by AI

Amazon Bedrock AgentCore now offers Runtime Instances, a managed EC2‑based compute option that extends session length, adds GPU access, and enables multiple agents to share a single host. This shift lets AI, cloud, and DevOps engineers build persistent, collaborative pipelines such as a multi‑day music production workflow while requiring new capacity‑provider and storage considerations.

Amazon Bedrock AgentCore introduced Runtime Instances, a managed EC2‑based option that replaces the short‑lived, single‑agent MicroVM model with persistent, multi‑day sessions, GPU support, and shared storage. Engineers building AI pipelines, platform teams managing compute, and SREs responsible for availability all need to understand how this new model reshapes deployment, scaling, and security practices.

How AgentCore Runtime Instances Change the Compute Model

Both MicroVMs and Runtime Instances expose the same AgentCore APIs, but the underlying compute differs. MicroVMs are fully managed, serverless containers that host a single agent per runtime and cap sessions at eight hours. Runtime Instances run on AWS‑managed EC2 instances, allowing a single instance to host multiple agents, extending sessions to fourteen days, and providing optional GPU access. Persistent storage is supplied via Amazon EBS, whereas MicroVM sessions are isolated and stateless. Pricing shifts from consumption‑based billing to the standard EC2 model, meaning existing Savings Plans or On‑Demand Capacity Reservations can be applied.

Architecting a Three‑Agent Music Production Pipeline

The example pipeline demonstrates three specialized agents sharing one GPU‑enabled Runtime Instance:

  1. Composition agent – packaged as a container image in Amazon ECR, it uses Claude Sonnet 4.6 to create a brief and then runs the ACE‑Step model on the instance GPU to generate a .wav file.
  2. Delivery agent – also a container in Amazon ECR, it reads the generated audio from a shared filesystem, measures it, asks Claude Sonnet 4.6 for an EQ/compression chain, applies the chain, and re‑measures to confirm target compliance.
  3. Compliance agent – delivered as a zip file from Amazon S3, it independently re‑measures the final track, validates the delivery targets, and screens the audio against a back‑catalog for harmonic similarity. If a conflict is detected, it invokes the composition agent for a replacement.

All three agents are bound to the same runtimeSessionId and a common capacity provider, ensuring they land on the same EC2 host and share an Amazon EBS volume. This co‑location enables direct file access without network hops and preserves state across days.

Operational Implications

  • Capacity providers must be defined to allocate the desired instance family and GPU type; they replace the on‑demand scaling model of MicroVMs.
  • Persistent EBS volumes require lifecycle management – snapshots for backup and proper encryption configuration.
  • Monitoring must cover both instance health (CPU, GPU, disk) and agent‑level metrics exposed via the AgentCore runtime APIs.
  • Cost estimation changes: instance uptime is billed continuously, so idle time should be minimized through session termination or instance recycling.

Security and Isolation Considerations

Runtime Instances keep the same session isolation guarantees as MicroVMs, but agents now share a filesystem. Practitioners should treat the shared EBS volume as a trusted intra‑instance boundary and enforce least‑privilege IAM policies for the agents’ access to Amazon ECR and Amazon S3. Because GPU drivers run on the host, any container‑level vulnerability could affect other agents on the same instance, suggesting the need for regular image scanning and runtime hardening.

Related CloudNinjas coverage: AWS.

What This Means For Practitioners

  • Evaluate whether your workflow requires multi‑day state or GPU acceleration; if so, migrate from MicroVMs to Runtime Instances.
  • Define capacity providers that match the GPU and storage profile of your agents, and bind them with a common runtimeSessionId to enable co‑location.
  • Implement robust EBS snapshot and encryption policies to protect persisted data across long sessions.
  • Update cost monitoring to account for continuous EC2 billing and consider ODCRs or Savings Plans to offset expense.
  • Apply container image scanning and least‑privilege IAM roles to mitigate the broader attack surface introduced by shared host resources.
Originally published atAWS Machine Learning Blog