Amazon Bedrock adds Z.ai's 753B-parameter GLM 5.3 for eligible enterprise customers
Amazon Bedrock now offers GLM 5.3, Z.ai's 753B-parameter mixture-of-experts model, through managed APIs with cross-Region inference, prompt caching and service tiers. Access is limited to eligible enterprise customers. Pricing and context window were not disclosed in AWS's announcement.
Amazon Bedrock has added GLM 5.3, a 753B-parameter mixture-of-experts (MoE) model from Z.ai (Zhipu AI), as a managed offering. According to AWS, access is available to eligible enterprise customers. GLM 5 arrived on Bedrock earlier this year.
This is an availability announcement. The model weights are already published on Hugging Face Hub, and the AWS post does not give a release date for GLM 5.3 itself.
What AWS confirmed
- Architecture: 753B-parameter MoE, as published on Hugging Face Hub. Active parameter count was not stated.
- Focus: Coding and long-horizon agentic tasks.
- APIs: The OpenAI-compatible Responses and Chat Completions APIs, plus Bedrock Invoke and Converse. AWS recommends the OpenAI-compatible APIs for new applications because they support a more complete feature set.
- Cross-Region inference profiles:
us.zai.glm-5.3(US) andglobal.zai.glm-5.3(Global). - Service tiers: Flex (lower cost, less time-sensitive), Standard (default), and Priority (lower latency, higher price).
- Prompt caching: Implicit caching is on by default. Explicit caching is recommended and is set through
prompt_cache_optionson Responses and Chat Completions requests. Eachprompt_cache_breakpointmust cover at least 1,024 tokens to be eligible.
AWS also notes improved feature parity between the OpenAI-compatible APIs and the Invoke and Converse APIs.
Performance claims
All benchmark figures below come from Z.ai, not independent testing.
- Z.ai claims competitive results on DeepSWE, Terminal Bench 3.0 and FrontierSWE. AWS's post gives no scores for these.
- Z.ai reports a 50% improvement over GLM 5.2 on its own internal coding benchmark.
- Z.ai reports a score of 84.5 on CyberGym at release and describes it as a leading result.
- No direct comparison to GLM 5 was provided. Z.ai says it updated its benchmark tests after the GLM 5.1 announcement because the gains were large, so scores across versions are not directly comparable.
Security-testing demo
AWS's walkthrough applies the model to authorized security testing of a user's own application with Strix, an open-source AI penetration-testing agent that runs code dynamically to find vulnerabilities. The setup requires Docker and Strix with the bedrock extra. Z.ai's cyber security claims are the stated reason for the choice, and AWS describes the model as a fit for defensive workflows.
Not disclosed
- Context window size
- Pricing per 1M input and output tokens (pricing not yet disclosed)
- Training data cutoff
- Active parameters per token
- Which customers count as "eligible"
What this means
For teams that want an open-weight frontier-scale coding model without running inference for a 753B MoE, Bedrock now offers one with caching, tiering and cross-Region routing. The OpenAI-compatible endpoint also lowers integration cost for applications already built on those SDKs.
The evidence for the model's quality is thin. The headline numbers are vendor-reported, the main coding comparison is against Z.ai's own internal benchmark, and there is no like-for-like comparison with GLM 5. Without published pricing or a context window, buyers cannot yet judge cost per task or how well the model handles repository-scale work. The restriction to eligible enterprise customers also limits who can run independent evaluations on Bedrock. Until third-party results appear, treat the CyberGym score and the 50% gain as claims to verify.
Related Articles
AWS adds managed Web Search to Claude Desktop via Bedrock AgentCore Gateway in three Regions
AWS published a walkthrough for connecting Claude Desktop on Amazon Bedrock to a managed, MCP-compatible Web Search capability through Amazon Bedrock AgentCore Gateway. According to AWS, the search is backed by an Amazon web index spanning tens of billions of documents, and query traffic stays within AWS infrastructure. Web Search is available in three AWS Regions; pricing is not disclosed in the post.
Claude Opus 5.5 and Sonnet 5.5 now on Amazon Bedrock in AWS GovCloud (US), with Claude Code support
Claude Opus 5.5 and Claude Sonnet 5.5 are available on Amazon Bedrock in AWS GovCloud (US) Regions. AWS published a setup guide for running Anthropic's Claude Code against them for regulated workloads, including ITAR. Pricing, context window and benchmark figures were not disclosed.
AWS ships aws-ai-ml skill so Kiro, Claude Code and Codex can benchmark SageMaker inference endpoints
Amazon SageMaker AI has released the aws-ai-ml skill, distributed through the Agent Toolkit for AWS. It lets MCP-compatible coding agents such as Kiro, Claude Code and Codex benchmark endpoints, recommend deployment configurations and generate executable SageMaker Python SDK v3 code.
AWS releases open-source MCP server to automate cross-account promotion of Amazon Quick agents
AWS has published the Quick Resource Migrator, a sample MCP server on Amazon Bedrock AgentCore that promotes Amazon Quick resources between AWS accounts in a single tool call. It is idempotent, never deletes from the target, and writes versioned S3 backups before every update.
Comments
Loading...