AWS ships aws-ai-ml skill so Kiro, Claude Code and Codex can benchmark SageMaker inference endpoints
Amazon SageMaker AI has released the aws-ai-ml skill, distributed through the Agent Toolkit for AWS. It lets MCP-compatible coding agents such as Kiro, Claude Code and Codex benchmark endpoints, recommend deployment configurations and generate executable SageMaker Python SDK v3 code.
Amazon SageMaker AI has released the aws-ai-ml skill, an agent skill that gives MCP-compatible coding agents, including Kiro, Claude Code and Codex, the ability to benchmark inference endpoints, recommend deployment configurations, compare performance runs and generate executable SageMaker Python SDK v3 code. It is distributed through the Agent Toolkit for AWS. This is a tooling release, not a new model.
What the skill does
According to AWS, the skill turns any agent that supports the Model Context Protocol (MCP) into a SageMaker AI inference optimization assistant. The user describes a goal in natural language, such as a latency target, a cost ceiling or a model to evaluate. The agent asks clarifying questions and produces code the user can review, modify and run in their own environment. AWS says the code is grounded in real benchmark and measured performance data.
The documented capabilities include:
- Benchmarking an existing endpoint. The agent generates a Python notebook that uses the
Workload.synthetic()andstart_benchmark()APIs from the SageMaker Python SDK. It first confirms the endpoint is safe to load-test, since benchmarking sends real traffic to a live endpoint. - Performance reporting. Reports cover throughput (requests per second, output tokens per second), latency (p50, p99, time-to-first-token, inter-token latency) and concurrency. AWS states these are measured values from real load, not estimates. The agent may also suggest improvements such as prefill decoding.
- Instance-type selection. Users can supply a fine-tuned model in Amazon S3 (by S3 URI), a SageMaker JumpStart model ID (the example given is
huggingface-reasoning-qwen3-8b) or a Hugging Face Hub model name. The source text cuts off while describing the Hugging Face Hub path, including gated-model handling.
Setup
There are two routes:
- Local install. This requires AWS CLI 2.35 or later and
uv. Runaws configure agent-toolkitto detect agents, install skills and configure the AWS MCP Server. Then runnpx skills add aws/agent-toolkit-for-aws/skills/aws-ai-ml. According to AWS, Kiro and Claude Code can also discover and load skills at runtime through the AWS MCP Server without a local install. - SageMaker Studio JupyterLab. Users select a pre-configured image that ships with the skill. The space must be set to Private, because skills only sync on private spaces. First boot takes 5–10 minutes. AWS recommends a fresh space, since a reused space with a locally modified skill may not pick up the image. Kiro authentication uses
kiro-cli login.
The skill generates code that runs under the user's own AWS credentials. Those credentials need permission to call SageMaker AI APIs, including creating endpoints and running benchmark and recommendation jobs. AWS says no additional IAM configuration is needed for the skill itself. AWS claims users can go from zero to a working conversation in 10 minutes.
Pricing for the skill itself is not stated in the source. Costs from endpoints and benchmark jobs the generated code creates would be incurred in the user's account.
What this means
The notable choice here is packaging. AWS is shipping SageMaker expertise as a skill for third-party agents, including Claude Code and Codex, rather than only through its own Kiro. That puts AWS inference tuning where developers already work.
The design also addresses a real problem. Picking an instance family and serving container for a given model is hard, and recommendations backed by measured throughput and latency are more useful than generic advice from a general-purpose coding agent. Because the output is SDK v3 code rather than an opaque UI action, teams can audit and version-control what the agent does.
The main caveat is operational. Load tests hit live endpoints, so the agent's safety check matters, and teams should still gate production benchmarking. AWS has not published comparative data on how much the skill improves cost or latency outcomes, so those gains remain unquantified.
Related Articles
AWS releases open-source MCP server to automate cross-account promotion of Amazon Quick agents
AWS has published the Quick Resource Migrator, a sample MCP server on Amazon Bedrock AgentCore that promotes Amazon Quick resources between AWS accounts in a single tool call. It is idempotent, never deletes from the target, and writes versioned S3 backups before every update.
AWS adds managed Web Search to Claude Desktop via Bedrock AgentCore Gateway in three Regions
AWS published a walkthrough for connecting Claude Desktop on Amazon Bedrock to a managed, MCP-compatible Web Search capability through Amazon Bedrock AgentCore Gateway. According to AWS, the search is backed by an Amazon web index spanning tens of billions of documents, and query traffic stays within AWS infrastructure. Web Search is available in three AWS Regions; pricing is not disclosed in the post.
OpenAI to watermark ChatGPT and Codex text in the EU with textGrain; API watermarking is opt-in worldwide
OpenAI will switch on invisible text watermarks called textGrain for ChatGPT and Codex users in the EU over the coming weeks. API watermarking will be opt-in worldwide, unlike Anthropic's mandatory approach for Claude. OpenAI's own data shows detection drops sharply when text is edited.
Claude Opus 5.5 and Sonnet 5.5 now on Amazon Bedrock in AWS GovCloud (US), with Claude Code support
Claude Opus 5.5 and Claude Sonnet 5.5 are available on Amazon Bedrock in AWS GovCloud (US) Regions. AWS published a setup guide for running Anthropic's Claude Code against them for regulated workloads, including ITAR. Pricing, context window and benchmark figures were not disclosed.
Comments
Loading...