Live
GitHub Rewrites Copilot Runtime in Rust via AI‑Guided Incremental MigrationECS auto‑repair for GPU and instance failures shifts remediation to the platformDecision Model API Converges on a Shared Schema – Implications for EngineersR2 dashboard now reports bandwidth per Cloudflare locationMinimum Viable Instrumentation adds gap detection to OllyGarden’s Rose AI agentWarehouse‑Native Extraction with Alteryx Live Query and BigQueryAI Agent Integration on Amazon Bedrock: Lessons from Postman's Production RolloutBedrock AgentCore Runtime Gains Speed, Pay‑As‑You‑Go, and New Model OptionsGitHub Rewrites Copilot Runtime in Rust via AI‑Guided Incremental MigrationECS auto‑repair for GPU and instance failures shifts remediation to the platformDecision Model API Converges on a Shared Schema – Implications for EngineersR2 dashboard now reports bandwidth per Cloudflare locationMinimum Viable Instrumentation adds gap detection to OllyGarden’s Rose AI agentWarehouse‑Native Extraction with Alteryx Live Query and BigQueryAI Agent Integration on Amazon Bedrock: Lessons from Postman's Production RolloutBedrock AgentCore Runtime Gains Speed, Pay‑As‑You‑Go, and New Model Options
AWS

Bedrock AgentCore Runtime Gains Speed, Pay‑As‑You‑Go, and New Model Options

AI SummaryPowered by AI

September 2026 brought a preview of Managed Agents with OpenAI support, a faster AgentCore runtime, token‑saving Strands tools, and new OpenAI models on Bedrock. These updates let engineers run agents more efficiently, enforce enterprise security controls, and choose models that fit cost and performance needs.

In September 2026 AWS released a set of updates that affect how AI agents are built, run, and cost‑optimized on Amazon Bedrock. The changes include a public preview of Managed Agents that support OpenAI models, a faster and more memory‑efficient AgentCore runtime, an open‑source Strands harness that cuts token usage, a lightweight Decider model for rapid option selection, and the general availability of new OpenAI models (Astra, Sol, Luna) with an ultrafast speed tier.

Faster Bedrock agent runtime and pay‑as‑you‑go usage

The AgentCore runtime now implements tighter memory management and reduces cold‑start latency for serverless agents. Agents can start with less allocated memory, and idle sessions automatically scale to zero, eliminating the need for pre‑provisioned capacity. The runtime runs in hardware‑isolated environments and charges only for actual usage, which can lower operational spend for bursty workloads.

Managed Agents with OpenAI models and enterprise controls

Amazon Bedrock Managed Agents entered public preview, allowing developers to use OpenAI models while keeping data inside AWS. The service reuses existing IAM permissions, logs all actions to CloudTrail, and supports durable sessions and built‑in human‑approval steps. These controls help teams meet audit and governance requirements without building custom scaffolding.

Token‑efficient open‑source tooling

Strands introduced a new harness that achieves the same accuracy as comparable frameworks while consuming 28 % fewer tokens. The harness can be instantiated with a single line of Python or TypeScript, and it provides built‑in context handling, prompt caching, and memory management, simplifying deployment across environments. Additionally, Strands Decider 2B offers a 2‑billion‑parameter model that selects among predefined options in roughly 115 ms on local hardware, making it suitable for routing, tool selection, and guardrail enforcement.

Expanded model portfolio on Bedrock

OpenAI’s Astra, Sol, and Luna models are now generally available on Bedrock. GPT‑6 Astra targets high‑complexity tasks such as deep reasoning, large document analysis, and code generation, supporting up to 1 million input tokens. An Ultrafast tier for Astra delivers up to six times faster inference, handling up to 300 tokens per second. GPT‑6.1 Sol provides near‑Astra capability for frequent coding and professional workloads.

Related CloudNinjas coverage: AWS.

What This Means For Practitioners

  • Evaluate whether the new AgentCore runtime can replace existing provisioned instances to reduce cold‑start latency and memory costs.
  • Leverage Managed Agents for OpenAI workloads when data residency and auditability are required, and map existing IAM policies to the service.
  • Consider adopting the Strands harness to lower token consumption, especially in high‑throughput pipelines.
  • Use Strands Decider 2B for fast, deterministic routing decisions instead of prompting large language models.
  • Match workload characteristics to the appropriate Bedrock model: Astra for complex, high‑token tasks; Ultrafast Astra for latency‑sensitive calls; Sol for routine coding or automation.
  • Monitor upcoming Bedrock announcements for further runtime optimizations or model additions that could affect cost and performance baselines.
Originally published atAWS Machine Learning Blog