AWS SageMaker HyperPod adds three-tier data capture, direct Hugging Face deployment, and NVMe caching for enterprise inf
Amazon SageMaker HyperPod has launched infrastructure updates for enterprise inference workloads. The platform now captures inference data at three points—endpoint, load balancer, and model pod—with configurable sampling and S3 storage. Teams can deploy models directly from Hugging Face Hub without pre-staging weights, with support for gated access across vLLM, TGI, and SGLang runtimes.
AWS SageMaker HyperPod adds three-tier data capture, direct Hugging Face deployment, and NVMe caching for enterprise inference
Amazon SageMaker HyperPod has launched a set of infrastructure capabilities designed for organizations running large model inference in production. The updates focus on observability, deployment flexibility, and performance optimization.
Three-tier data capture system
HyperPod now captures inference request and response data at three independent points in the inference path:
Tier 1 - SageMaker endpoint: Captures full input/output payloads at the SageMaker Runtime API boundary. Required for SageMaker Model Monitor compatibility. Data stored at {s3Uri}/{hash}/sme/.
Tier 2 - Application Load Balancer: Enables ALB access logs capturing request metadata including client IPs, request paths, and latencies. Data stored at {s3Uri}/{hash}/alb/.
Tier 3 - Model pod: Captures full inference payloads at the container level with configurable sampling (0-100%), buffer batch size, flush intervals, and payload size limits up to specified KB. Works without endpoint registration. Data stored at {s3Uri}/{hash}/pod/.
Each tier can be enabled independently through declarative CRD configuration. Teams can set sampling percentages per tier, specify KMS encryption keys, and define S3 storage locations. If no S3 URI is specified, HyperPod defaults to the TLS certificate bucket with a /data-capture/ prefix.
According to AWS, the same deployment generates the same S3 prefix across multiple CRD submissions, consolidating capture artifacts from a single deployment into one subfolder.
Direct Hugging Face Hub deployment
HyperPod now supports deploying models directly from Hugging Face Hub without pre-staging weights to S3 or FSx storage. The integration includes:
- Gated model access through
tokenSecretRefstored in Kubernetes Secrets - Revision pinning via
commitSHAspecification - Token isolation across deployments
- Compatibility with vLLM, TGI, and SGLang inference runtimes
The Inference Operator emits Kubernetes events tracking deployment status and failures. Required fields include a valid modelId, a Kubernetes Secret containing the Hugging Face API token for gated models, and a GPU-enabled worker node with volume mounts for downloaded weights.
NVMe caching and DNS management
The platform now loads model weights from node-local NVMe storage to reduce cold-start latency, with automatic fallback to cloud storage when local storage is unavailable. Pricing for NVMe storage not disclosed.
HyperPod automatically manages custom domain DNS records through Route 53 integration. Infrastructure teams can set granular pod-level IAM permissions for security boundaries.
IAM requirements
To enable data capture on existing clusters, teams must add S3 PutObject permissions to the Inference Operator Execution Role with resource scope arn:aws:s3:::hyperpod-tls*/data-capture/*. Customer-managed KMS keys require additional kms:Decrypt and kms:GenerateDataKey permissions.
AWS recommends starting with lower sampling percentages in production, using maxPayloadSizeKB limits to control storage costs, and placing sensitive data in POST request bodies rather than query parameters since ALB logs (Tier 2) capture URLs and query strings.
What this means
These updates address three operational bottlenecks in production AI infrastructure: observability without custom logging pipelines, deployment friction from weight staging requirements, and cold-start latency from remote storage reads. The three-tier capture system is particularly relevant for regulated industries requiring audit trails at multiple infrastructure layers. Direct Hugging Face integration removes a manual step that previously required teams to download, version, and upload multi-gigabyte model files before deployment. The NVMe caching targets the cold-start problem that affects autoscaling inference workloads, though AWS has not disclosed performance benchmarks comparing NVMe versus S3/FSx load times.
Related Articles
Amazon Bedrock adds Z.ai's 753B-parameter GLM 5.3 for eligible enterprise customers
Amazon Bedrock now offers GLM 5.3, Z.ai's 753B-parameter mixture-of-experts model, through managed APIs with cross-Region inference, prompt caching and service tiers. Access is limited to eligible enterprise customers. Pricing and context window were not disclosed in AWS's announcement.
AWS ships aws-ai-ml skill so Kiro, Claude Code and Codex can benchmark SageMaker inference endpoints
Amazon SageMaker AI has released the aws-ai-ml skill, distributed through the Agent Toolkit for AWS. It lets MCP-compatible coding agents such as Kiro, Claude Code and Codex benchmark endpoints, recommend deployment configurations and generate executable SageMaker Python SDK v3 code.
Common Sense Media rates ChatGPT for Teens an 'unacceptable risk,' citing failures on 3 of 5 Red Lines
Common Sense Media has labeled OpenAI's ChatGPT for Teens an "unacceptable risk," saying it failed three of five severe-harm Red Lines and kept using engagement cues during crisis conversations. OpenAI disputes the testing methodology, saying testing may have ended before parental controls were fully active.
Microsoft's Copilot gets access to local Windows files and OS-level actions under 'Hybrid Intelligence'
Microsoft announced an upgrade to Copilot at its Windows and Surface event that gives the assistant access to local files and the ability to take actions across Windows. The company calls the underlying approach "Hybrid Intelligence," which combines local and cloud AI models. Pricing, model details, and availability were not disclosed in the available reporting.
Comments
Loading...