Amazon SageMaker AI has introduced the aws-ai-ml skill through the Agent Toolkit for AWS, enabling MCP-compatible coding agents to generate executable SageMaker Python SDK v3 code for inference optimization workflows. For engineers, that means a natural-language request can yield code to review, modify and run under their own AWS credentials rather than an opaque automated deployment action.

AWS describes the skill as an integration for coding agents including Kiro, Claude Code and Codex, not a new model-serving product. It is designed to translate goals such as meeting a performance target, testing a deployed model or staying within a cost constraint into SageMaker benchmarking and recommendation workflows. The agent asks for missing inputs instead of guessing, according to AWS, and exposes its work as code.

Benchmarking a live endpoint

For a model already hosted on a SageMaker AI endpoint, the skill can generate a Python notebook that uses the SageMaker SDK’s Workload.synthetic() and start_benchmark() APIs to run a load test. AWS says the resulting report includes request and output-token throughput, p50 and p99 latency, time to first token, inter-token latency and supported concurrency.

Those benchmark values come from real traffic sent to real infrastructure, AWS says, rather than from estimates. That also creates a material operational constraint: before starting a benchmark, the agent is meant to confirm that the live endpoint is safe to load-test. It can recommend improvement mechanisms such as prefill decoding after a run, but users retain responsibility for the traffic and resources created.

Evaluating candidate deployments

The skill can also generate code to test a model against candidate instance types and configurations, then present ranked options using throughput, latency percentiles, time to first token and concurrency. AWS says this workflow supports custom or fine-tuned models stored in Amazon S3, public foundation models in SageMaker JumpStart, and models hosted on the Hugging Face Hub.

For gated Hugging Face models, the agent is intended to surface license terms and ask the user to accept them and provide a Hugging Face token. The skill can also compare two named benchmark jobs, calculating percentage deltas for throughput, latency and time-to-first-token metrics. If a requested run does not exist, AWS says the agent can offer to run it before attempting the comparison.

Published example does not isolate model choice

AWS published an example based on a 512/256-token workload at concurrency four. In that comparison, Qwen3-8B recorded 271.2 output tokens per second, compared with 188.2 for Qwen3-1.7B, a 44.1% difference. Request throughput was 1.08 requests per second versus 0.736, and request latency was 3,658 ms versus 5,382 ms. Inter-token latency was 14 ms for Qwen3-8B and 20.9 ms for Qwen3-1.7B.

The figures should not be read as a clean comparison of model sizes. Qwen3-8B ran on a four-GPU ml.g5.12xlarge using four A10G GPUs, while Qwen3-1.7B used one L4 GPU on an ml.g6.4xlarge. AWS says the comparison reflects roughly four times as much compute. The smaller model had the faster time to first token, at 67.5 ms versus 166.3 ms.

Setup paths and user controls

The skill can be installed locally through the Agent Toolkit for AWS or used in an Amazon SageMaker Studio JupyterLab space configured with the skill. AWS says local setup requires AWS CLI 2.35 or later and uv. AWS credentials need permission to call SageMaker AI APIs for endpoint creation and benchmark or recommendation jobs; the generated code runs with those credentials. AWS says Kiro and Claude Code can discover skills at runtime through the AWS MCP Server.

In Studio, skills synchronize only in private JupyterLab spaces. AWS says an initially configured space can take five to 10 minutes to boot and recommends a fresh space, since a reused space with a locally modified skill may not receive the pre-configured version. The agent is also intended to explain when a request is outside its scope, such as directly deploying a model, and provide deployment configuration instead.

Costs remain with the user

The skill does not remove the need to manage SageMaker resources. AWS advises users to delete endpoints created during benchmarks or recommendations, stop or delete JupyterLab spaces, and remove S3 objects generated by benchmark and recommendation jobs to avoid ongoing charges.

Source: AWS Machine Learning Blog.

Definition. The aws-ai-ml skill is an Agent Toolkit for AWS integration that turns natural-language inference optimization goals into SageMaker code users can review, modify and run with their own credentials.

MetricQwen3-8B vs. Qwen3-1.7B
Output tokens per second271.2 vs. 188.2
Request throughput1.08 vs. 0.736 requests per second
Request latency3,658 ms vs. 5,382 ms
Inter-token latency14 ms vs. 20.9 ms
Time to first token166.3 ms vs. 67.5 ms

Key takeaways

  • The skill is an integration for coding agents, not a new model-serving product.
  • Live endpoint benchmarks send real traffic to real infrastructure, so users must confirm an endpoint is safe to load-test.
  • Generated workflows can rank candidate configurations by throughput, latency, time to first token and concurrency.
  • AWS’s published Qwen example used different hardware configurations and does not isolate model-size effects.
  • Users remain responsible for credentials, created resources and cleanup costs.

FAQ

What does the aws-ai-ml skill generate?

It generates executable SageMaker Python SDK v3 code for inference benchmarking, deployment recommendations and benchmark comparisons.

Can the skill benchmark a live SageMaker endpoint?

Yes. It can generate a notebook using Workload.synthetic() and start_benchmark() to load-test an existing endpoint after confirming it is safe to do so.

Which benchmark metrics can the workflow report?

AWS says reports can include request and output-token throughput, p50 and p99 latency, time to first token, inter-token latency and supported concurrency.

Does the published Qwen comparison show model-size performance alone?

No. Qwen3-8B and Qwen3-1.7B used different instance types and roughly four times different compute, so the figures do not isolate model size.

Who is responsible for costs created by the skill?

The user remains responsible for traffic, SageMaker resources and cleanup, including endpoints, JupyterLab spaces and generated S3 objects.

Sources