Amazon Bedrock now exposes Anthropic’s Claude Opus 5 and Claude Sonnet 5 models through the bedrock-runtime endpoint with true in‑region inference in the Asia Pacific (Seoul) ap-northeast-2 region and the Asia Pacific (Singapore) ap-southeast-1 region. The change means that inference requests, prompts, and generated outputs are processed entirely inside the selected AWS region, never leaving its network boundary.
Supported models and regions
The following model identifiers are available for in‑region calls:
anthropic.claude-opus-5– Seoul onlyanthropic.claude-sonnet-5– Seoul and Singapore
Both models can be invoked via the standard Bedrock APIs: the Anthropic Messages API, the generic InvokeModel API, and the conversational Converse API. Guardrails and intelligent prompt routing remain usable because they are part of the Bedrock runtime.
How in‑region inference works
When a request is sent to a specific region, Bedrock does not route it through a cross‑region layer. The request is handled by the compute resources in that region alone, and all data—including the input prompt and the model’s response—remains there for the full request lifecycle. Because there is no inter‑region hop, latency is limited to the region’s internal network, but throughput is capped by the region’s capacity and its per‑region service quotas. Billing follows the standard on‑demand rates for the region that receives the request, and monitoring data (CloudWatch metrics, CloudTrail logs) is scoped to the same region.
Getting started
Practitioners can experiment without code via the Bedrock console’s Playground. After selecting the appropriate region, the model is chosen by searching for its identifier (e.g., anthropic.claude-opus-5) and setting the inference type to “On‑Demand”. A prompt can then be submitted and the response displayed instantly.
For programmatic access, the same model IDs are passed to the bedrock-runtime endpoint. The request payload follows the format defined by the Anthropic Messages API or the generic InvokeModel schema, depending on the integration style. The endpoint URL includes the region code, ensuring the call is routed to the intended in‑region service.
Operational implications
- Capacity planning: Since each region enforces its own quota, teams must monitor regional usage and request quota increases if needed.
- Compliance handling: Data residency requirements for regulated sectors (financial services, healthcare, public sector) can be satisfied by keeping all model traffic inside the required geography.
- Observability: CloudWatch dashboards and CloudTrail audit trails will only contain entries from the region that processed the request, simplifying regional log aggregation but requiring separate setups for each region.
- Cost awareness: On‑demand pricing is applied per region, so cost estimates should reflect the region‑specific rates rather than a global average.
Related CloudNinjas coverage: AWS.
What This Means For Practitioners
Engineers can now integrate Anthropic’s most capable Claude models into workloads that must remain within South Korea or Singapore without adding a separate data‑transfer layer. The immediate action is to update automation scripts to target the correct ap-northeast-2 or ap-southeast-1 bedrock-runtime endpoint and to verify that regional quotas and monitoring pipelines are in place. Teams should also evaluate whether the regional capacity meets their latency and throughput targets, and request quota adjustments early in the project lifecycle.



