product updateAmazon Web Services

Amazon Bedrock adds Z.ai's 753B-parameter GLM 5.3 for eligible enterprise customers

TL;DR

Amazon Bedrock now offers GLM 5.3, Z.ai's 753B-parameter mixture-of-experts model, through managed APIs with cross-Region inference, prompt caching and service tiers. Access is limited to eligible enterprise customers. Pricing and context window were not disclosed in AWS's announcement.

3 min read
0

Amazon Bedrock has added GLM 5.3, a 753B-parameter mixture-of-experts (MoE) model from Z.ai (Zhipu AI), as a managed offering. According to AWS, access is available to eligible enterprise customers. GLM 5 arrived on Bedrock earlier this year.

This is an availability announcement. The model weights are already published on Hugging Face Hub, and the AWS post does not give a release date for GLM 5.3 itself.

What AWS confirmed

  • Architecture: 753B-parameter MoE, as published on Hugging Face Hub. Active parameter count was not stated.
  • Focus: Coding and long-horizon agentic tasks.
  • APIs: The OpenAI-compatible Responses and Chat Completions APIs, plus Bedrock Invoke and Converse. AWS recommends the OpenAI-compatible APIs for new applications because they support a more complete feature set.
  • Cross-Region inference profiles: us.zai.glm-5.3 (US) and global.zai.glm-5.3 (Global).
  • Service tiers: Flex (lower cost, less time-sensitive), Standard (default), and Priority (lower latency, higher price).
  • Prompt caching: Implicit caching is on by default. Explicit caching is recommended and is set through prompt_cache_options on Responses and Chat Completions requests. Each prompt_cache_breakpoint must cover at least 1,024 tokens to be eligible.

AWS also notes improved feature parity between the OpenAI-compatible APIs and the Invoke and Converse APIs.

Performance claims

All benchmark figures below come from Z.ai, not independent testing.

  • Z.ai claims competitive results on DeepSWE, Terminal Bench 3.0 and FrontierSWE. AWS's post gives no scores for these.
  • Z.ai reports a 50% improvement over GLM 5.2 on its own internal coding benchmark.
  • Z.ai reports a score of 84.5 on CyberGym at release and describes it as a leading result.
  • No direct comparison to GLM 5 was provided. Z.ai says it updated its benchmark tests after the GLM 5.1 announcement because the gains were large, so scores across versions are not directly comparable.

Security-testing demo

AWS's walkthrough applies the model to authorized security testing of a user's own application with Strix, an open-source AI penetration-testing agent that runs code dynamically to find vulnerabilities. The setup requires Docker and Strix with the bedrock extra. Z.ai's cyber security claims are the stated reason for the choice, and AWS describes the model as a fit for defensive workflows.

Not disclosed

  • Context window size
  • Pricing per 1M input and output tokens (pricing not yet disclosed)
  • Training data cutoff
  • Active parameters per token
  • Which customers count as "eligible"

What this means

For teams that want an open-weight frontier-scale coding model without running inference for a 753B MoE, Bedrock now offers one with caching, tiering and cross-Region routing. The OpenAI-compatible endpoint also lowers integration cost for applications already built on those SDKs.

The evidence for the model's quality is thin. The headline numbers are vendor-reported, the main coding comparison is against Z.ai's own internal benchmark, and there is no like-for-like comparison with GLM 5. Without published pricing or a context window, buyers cannot yet judge cost per task or how well the model handles repository-scale work. The restriction to eligible enterprise customers also limits who can run independent evaluations on Bedrock. Until third-party results appear, treat the CyberGym score and the 50% gain as claims to verify.

Related Articles

product update

AWS adds managed Web Search to Claude Desktop via Bedrock AgentCore Gateway in three Regions

AWS published a walkthrough for connecting Claude Desktop on Amazon Bedrock to a managed, MCP-compatible Web Search capability through Amazon Bedrock AgentCore Gateway. According to AWS, the search is backed by an Amazon web index spanning tens of billions of documents, and query traffic stays within AWS infrastructure. Web Search is available in three AWS Regions; pricing is not disclosed in the post.

product update

Claude Opus 5.5 and Sonnet 5.5 now on Amazon Bedrock in AWS GovCloud (US), with Claude Code support

Claude Opus 5.5 and Claude Sonnet 5.5 are available on Amazon Bedrock in AWS GovCloud (US) Regions. AWS published a setup guide for running Anthropic's Claude Code against them for regulated workloads, including ITAR. Pricing, context window and benchmark figures were not disclosed.

product update

AWS ships aws-ai-ml skill so Kiro, Claude Code and Codex can benchmark SageMaker inference endpoints

Amazon SageMaker AI has released the aws-ai-ml skill, distributed through the Agent Toolkit for AWS. It lets MCP-compatible coding agents such as Kiro, Claude Code and Codex benchmark endpoints, recommend deployment configurations and generate executable SageMaker Python SDK v3 code.

product update

AWS releases open-source MCP server to automate cross-account promotion of Amazon Quick agents

AWS has published the Quick Resource Migrator, a sample MCP server on Amazon Bedrock AgentCore that promotes Amazon Quick resources between AWS accounts in a single tool call. It is idempotent, never deletes from the target, and writes versioned S3 backups before every update.

Comments

Loading...