GLM 5.3
Z.AI — GLM 5.3
Model Details
GLM 5.3 is Z.AI's flagship open weight text generation model for complex software engineering and agentic tasks. It uses a Mixture-of-Experts architecture with 744B total parameters and approximately 40B active per token, and is built on the same base model as GLM 5.2 with improvements driven by scaled post-training. The model provides a 1M-token context window with up to 128K output tokens, and supports selectable reasoning effort levels so you can trade latency and token consumption against task performance. Typical uses include agentic coding and autonomous software engineering, terminal and CLI task execution, repository-scale code generation and refactoring, long-horizon multi-step engineering tasks spanning many tool calls, infrastructure diagnosis, and long-context analysis of large codebases. It supports function calling, structured JSON output, and response streaming including streamed reasoning content. For more information about model development and performance, see the model card
Note
Access to this model is limited to eligible customers. For more information, contact your AWS account team.
Model launch date: October 5, 2026
EOL no sooner than: Not Applicable, at least 45 day EOL Notice will be provided
Legacy period: at least 45 days
Model lifecycle policy: Bedrock Model Lifecycle
Model EOL date: N/A
Model lifecycle: Active
Context window: 1M tokens
Max output tokens: 128K
End User License Agreements and Terms of Use: View
| Input Modalities | Output Modalities | APIs supported | Endpoints supported |
|---|---|---|---|
Responses | bedrock-runtime | ||
Chat Completions | bedrock-mantle | ||
Converse | |||
Invoke | |||
Note
GLM 5.3 is available through cross-Region inference only. Requests must use the US geographic inference profile (us.zai.glm-5.3) or the global inference profile (global.zai.glm-5.3). In-Region on-demand inference is not supported.
Capabilities and Features
Bedrock Features
Features supported using bedrock-runtime endpoint
Explicit prompt caching using bedrock-runtime endpoint
For more information, see Prompt caching for faster model inference.
| Explicit Prompt Caching supported | Min tokens per cache checkpoint | Cache retention (TTL) |
|---|---|---|
| Yes | 1,024 | At least 30 minutes |
Note
By default, GLM 5.3 supports implicit (automatic) prompt caching. Configuring explicit cache controls can improve your cache hit rate, and therefore reduce latency and cost, so we recommend using explicit prompt caching. See the prompt caching guide for more details.
Pricing
For pricing information, see the Amazon Bedrock Pricing
Programmatic Access
Use the following model IDs and endpoint URLs to access this model programmatically. GLM 5.3 is available through cross-Region inference only, so you must use the geo or global inference ID rather than the base model ID. For more information about the available APIs and endpoints, see APIs supported and Endpoints supported.
| Endpoint | Model ID | In-Region endpoint URL | Geo inference ID | Global inference ID |
|---|---|---|---|---|
bedrock-runtime |
zai.glm-5.3 |
https://bedrock-runtime.{region}.amazonaws.com |
us.zai.glm-5.3 |
global.zai.glm-5.3 |
For example, if region is us-east-1 (N. Virginia), then the bedrock-runtime endpoint URL will be "https://bedrock-runtime.us-east-1.amazonaws.com".
Service Tiers
Amazon Bedrock offers multiple service tiers to match your workload requirements. Standard provides pay-per-token access with no commitment (set "service_tier": "default" or omit the field). Priority delivers the fastest response times for a price premium (set "service_tier": "priority"). Flex provides lower-cost access for flexible, non-time-sensitive workloads (set "service_tier": "flex"). Reserved provides dedicated throughput with a term commitment for predictable workloads; it is set at the account level rather than per request (contact your AWS account team to enable). For more information, see service tiers.
| Standard | Priority | Flex | Reserved |
|---|---|---|---|
Regional Availability
Regional availability at a glance
Amazon Bedrock offers three inference options: In-Region keeps requests within a single Region for strict compliance, Geo Cross-Region routes across Regions within a geography (such as US, EU, and APAC) while respecting data residency, and Global Cross-Region routes anywhere worldwide when there are no residency constraints. Refer to the Regional availability by models page for more details.
GLM 5.3 supports Geo Cross-Region inference from the US Regions and Global Cross-Region inference from the Regions listed in the following table. In-Region inference is not supported.
| Region | In-Region | Geo | Global |
|---|---|---|---|
us-east-1 (N. Virginia) | |||
us-east-2 (Ohio) | |||
us-west-1 (N. California) | |||
us-west-2 (Oregon) | |||
ca-central-1 (Canada) | |||
ca-west-1 (Calgary) | |||
eu-central-1 (Frankfurt) | |||
eu-central-2 (Zurich) | |||
eu-north-1 (Stockholm) | |||
eu-south-1 (Milan) | |||
eu-south-2 (Spain) | |||
eu-west-1 (Ireland) | |||
eu-west-2 (London) | |||
eu-west-3 (Paris) | |||
ap-east-2 (Taipei) | |||
ap-northeast-1 (Tokyo) | |||
ap-northeast-2 (Seoul) | |||
ap-northeast-3 (Osaka) | |||
ap-south-1 (Mumbai) | |||
ap-south-2 (Hyderabad) | |||
ap-southeast-1 (Singapore) | |||
ap-southeast-2 (Sydney) | |||
ap-southeast-3 (Jakarta) | |||
ap-southeast-4 (Melbourne) | |||
ap-southeast-5 (Malaysia) | |||
ap-southeast-6 (New Zealand) | |||
ap-southeast-7 (Thailand) | |||
il-central-1 (Tel Aviv) | |||
af-south-1 (Cape Town) | |||
mx-central-1 (Mexico) | |||
sa-east-1 (São Paulo) |
Quotas and Limits
Your AWS account has default quotas to maintain the performance of the service and to ensure appropriate usage of Amazon Bedrock. The default quotas assigned to an account might be updated depending on regional factors, payment history, fraudulent usage, and/or approval of a quota increase request. For more information, see Quotas for Amazon Bedrock documentation and see the limits for the model.
Sample Code
Step 1 - AWS Account: If you have an AWS account already, skip this step. If you are new to AWS, sign up for an AWS account
Step 2 - API key: Go to the Amazon Bedrock console
Step 3 - Get the SDK: To use this getting started guide, you must have Python already installed. Then install the relevant software depending on the APIs you are using.
Step 4 - Set environment variables: Configure your environment to use the API key for authentication.
Step 5 - Run your first inference request: Save the file as bedrock-first-request.py