Amazon SageMaker

3 articles tagged with Amazon SageMaker

September 25, 2026
product updateAmazon Web Services+1

AWS Brings Alibaba's Qwen3-TTS Voice Cloning Model to SageMaker Real-Time Endpoints

AWS published a deployment guide for running Alibaba's Qwen3-TTS-12Hz-1.7B-Base voice cloning model as a real-time SageMaker inference endpoint. The model clones a speaker's voice from a short audio clip and generates speech in 10 languages, including cross-lingual cloning, without retraining.

September 18, 2026
product updateAmazon Web Services

AWS Ships 13 SageMaker Inference Features in 2026, Cutting Startup Latency 51% and GPT-OSS-20B Throughput 2x

Amazon rolled out 13 new SageMaker AI inference capabilities in 2026 across managed endpoints and HyperPod Inference, spanning automated benchmarking, instance-pool fallback, OpenAI-compatible APIs, and container caching. AWS claims container caching cut endpoint startup latency by 51% and an inference-recommendation feature doubled GPT-OSS-20B throughput at equal latency.

August 17, 2026
model releaseNVIDIA+1

NVIDIA Nemotron 3.5 Lightning Arrives on Amazon SageMaker JumpStart, Targets High-Volume Agentic Workloads

NVIDIA's Nemotron 3.5 Lightning, a 30B-parameter hybrid Mixture-of-Experts model with only 3B active parameters, is now available for one-click deployment on Amazon SageMaker JumpStart. NVIDIA claims up to 4x higher throughput and 30% faster task completion for high-volume agentic workloads compared to larger frontier models.