In September 2026 AWS released a set of updates that affect how AI agents are built, run, and cost‑optimized on Amazon Bedrock. The changes include a public preview of Managed Agents that support OpenAI models, a faster and more memory‑efficient AgentCore runtime, an open‑source Strands harness that cuts token usage, a lightweight Decider model for rapid option selection, and the general availability of new OpenAI models (Astra, Sol, Luna) with an ultrafast speed tier.
Faster Bedrock agent runtime and pay‑as‑you‑go usage
The AgentCore runtime now implements tighter memory management and reduces cold‑start latency for serverless agents. Agents can start with less allocated memory, and idle sessions automatically scale to zero, eliminating the need for pre‑provisioned capacity. The runtime runs in hardware‑isolated environments and charges only for actual usage, which can lower operational spend for bursty workloads.
Managed Agents with OpenAI models and enterprise controls
Amazon Bedrock Managed Agents entered public preview, allowing developers to use OpenAI models while keeping data inside AWS. The service reuses existing IAM permissions, logs all actions to CloudTrail, and supports durable sessions and built‑in human‑approval steps. These controls help teams meet audit and governance requirements without building custom scaffolding.
Token‑efficient open‑source tooling
Strands introduced a new harness that achieves the same accuracy as comparable frameworks while consuming 28 % fewer tokens. The harness can be instantiated with a single line of Python or TypeScript, and it provides built‑in context handling, prompt caching, and memory management, simplifying deployment across environments. Additionally, Strands Decider 2B offers a 2‑billion‑parameter model that selects among predefined options in roughly 115 ms on local hardware, making it suitable for routing, tool selection, and guardrail enforcement.
Expanded model portfolio on Bedrock
OpenAI’s Astra, Sol, and Luna models are now generally available on Bedrock. GPT‑6 Astra targets high‑complexity tasks such as deep reasoning, large document analysis, and code generation, supporting up to 1 million input tokens. An Ultrafast tier for Astra delivers up to six times faster inference, handling up to 300 tokens per second. GPT‑6.1 Sol provides near‑Astra capability for frequent coding and professional workloads.
Related CloudNinjas coverage: AWS.
What This Means For Practitioners
- Evaluate whether the new AgentCore runtime can replace existing provisioned instances to reduce cold‑start latency and memory costs.
- Leverage Managed Agents for OpenAI workloads when data residency and auditability are required, and map existing IAM policies to the service.
- Consider adopting the Strands harness to lower token consumption, especially in high‑throughput pipelines.
- Use Strands Decider 2B for fast, deterministic routing decisions instead of prompting large language models.
- Match workload characteristics to the appropriate Bedrock model: Astra for complex, high‑token tasks; Ultrafast Astra for latency‑sensitive calls; Sol for routine coding or automation.
- Monitor upcoming Bedrock announcements for further runtime optimizations or model additions that could affect cost and performance baselines.


