Amazon SageMaker AI has introduced the aws-ai-ml skill for the Agent Toolkit for AWS, allowing any coding agent that implements the Model Context Protocol (MCP) to act as a SageMaker inference‑optimization assistant. The skill can benchmark endpoints, suggest deployment configurations, compare runs, and emit ready‑to‑run SageMaker Python SDK v3 snippets, all from natural‑language prompts.
What the skill does
The aws-ai-ml skill embeds three core capabilities into a coding agent:
- Benchmarking: The agent can invoke SageMaker AI APIs to run performance tests against a model endpoint and retrieve measured latency and throughput.
- Recommendation: Based on the benchmark data and user‑provided constraints (cost ceiling, latency target, instance family preferences), the agent proposes a concrete deployment configuration.
- Code generation: The agent emits SageMaker Python SDK v3 code that creates the recommended endpoint, sets scaling policies, and optionally configures VPC isolation.
All interactions remain visible as generated code, giving engineers the ability to review, edit, and execute the output in their own environment.
Getting the skill into your workflow
Two deployment paths are supported.
Option A – Attach to any MCP‑compatible agent (Kiro, Claude Code, Codex, etc.)
- Install the Agent Toolkit for AWS, which requires AWS CLI 2.35+ and
uv.aws configure agent-toolkit - Add the
aws-ai-mlskill.npx skills add aws/agent-toolkit-for-aws/skills/aws-ai-ml - Verify the skill appears in the agent’s skill list and start a conversation, e.g., “Generate a SageMaker endpoint for a 2‑GB LLM that must stay under $0.10 per hour.”
Option B – Use inside SageMaker Studio
Open a Studio JupyterLab space, launch a pre‑configured image that includes the Agent Toolkit, and repeat steps 1–3 from the local workflow. The skill runs in the same AWS account and region as the Studio environment.
Prerequisite permissions include the ability to call SageMaker AI endpoint creation, benchmark jobs, and any related IAM actions. No additional IAM resources are created by the skill itself.
Architectural and operational considerations
From an architecture perspective, the skill is a client‑side plug‑in; it does not introduce a new AWS service. The agent still communicates directly with SageMaker APIs using the caller’s credentials, so existing network topology, VPC endpoints, and data‑transfer controls remain unchanged.
Operationally, the generated code can be incorporated into CI/CD pipelines, but teams should treat the benchmark runs as workload‑specific tests that may incur charges. Because the skill can suggest instance families and scaling policies, engineers should validate cost estimates against their budgeting tools before production rollout.
Security implications are limited to the credential scope required to invoke SageMaker APIs. Since the skill does not store secrets, the primary risk is granting overly permissive IAM policies. Practitioners should follow the principle of least privilege, granting only the SageMaker actions needed for endpoint creation and benchmarking.
Related CloudNinjas coverage: AWS.
What This Means For Practitioners
Adopting the aws-ai-ml skill gives engineers a programmable bridge between high‑level performance goals and concrete SageMaker deployment artifacts. Teams can reduce manual benchmarking effort, obtain data‑driven configuration suggestions, and keep the entire decision flow in code that is auditable and version‑controlled. The next steps are to pilot the skill with a non‑critical model, verify the cost of benchmark jobs, and ensure IAM policies are tightly scoped before scaling the approach to production workloads.

