View a markdown version of this page

GLM 5.3 - Amazon Bedrock

GLM 5.3

Z.AI logo. Z.AI — GLM 5.3

Model Details

GLM 5.3 is Z.AI's flagship open weight text generation model for complex software engineering and agentic tasks. It uses a Mixture-of-Experts architecture with 744B total parameters and approximately 40B active per token, and is built on the same base model as GLM 5.2 with improvements driven by scaled post-training. The model provides a 1M-token context window with up to 128K output tokens, and supports selectable reasoning effort levels so you can trade latency and token consumption against task performance. Typical uses include agentic coding and autonomous software engineering, terminal and CLI task execution, repository-scale code generation and refactoring, long-horizon multi-step engineering tasks spanning many tool calls, infrastructure diagnosis, and long-context analysis of large codebases. It supports function calling, structured JSON output, and response streaming including streamed reasoning content. For more information about model development and performance, see the model card.

Note

Access to this model is limited to eligible customers. For more information, contact your AWS account team.

  • Model launch date: October 5, 2026

  • EOL no sooner than: Not Applicable, at least 45 day EOL Notice will be provided

  • Legacy period: at least 45 days

  • Model lifecycle policy: Bedrock Model Lifecycle

  • Model EOL date: N/A

  • Model lifecycle: Active

  • Context window: 1M tokens

  • Max output tokens: 128K

  • End User License Agreements and Terms of Use: View

Input Modalities Output Modalities APIs supported Endpoints supported
Red circle with white X icon indicating error, cancel, or close action. AudioRed circle with white X icon indicating error, cancel, or close action. EmbeddingGreen circle with white checkmark icon. ResponsesGreen circle with white checkmark icon. bedrock-runtime
Red circle with white X icon indicating error, cancel, or close action. ImageRed circle with white X icon indicating error, cancel, or close action. ImageGreen circle with white checkmark icon. Chat CompletionsRed circle with white X icon indicating error, cancel, or close action. bedrock-mantle
Red circle with white X icon indicating error, cancel, or close action. SpeechRed circle with white X icon indicating error, cancel, or close action. SpeechGreen circle with white checkmark icon. Converse
Green circle with white checkmark icon. TextGreen circle with white checkmark icon. TextGreen circle with white checkmark icon. Invoke
Red circle with white X icon indicating error, cancel, or close action. VideoRed circle with white X icon indicating error, cancel, or close action. Video
Note

GLM 5.3 is available through cross-Region inference only. Requests must use the US geographic inference profile (us.zai.glm-5.3) or the global inference profile (global.zai.glm-5.3). In-Region on-demand inference is not supported.

Capabilities and Features

Bedrock Features

Features supported using bedrock-runtime endpoint

Explicit prompt caching using bedrock-runtime endpoint

For more information, see Prompt caching for faster model inference.

Explicit Prompt Caching supported Min tokens per cache checkpoint Cache retention (TTL)
Yes 1,024 At least 30 minutes
Note

By default, GLM 5.3 supports implicit (automatic) prompt caching. Configuring explicit cache controls can improve your cache hit rate, and therefore reduce latency and cost, so we recommend using explicit prompt caching. See the prompt caching guide for more details.

Pricing

For pricing information, see the Amazon Bedrock Pricing page.

Programmatic Access

Use the following model IDs and endpoint URLs to access this model programmatically. GLM 5.3 is available through cross-Region inference only, so you must use the geo or global inference ID rather than the base model ID. For more information about the available APIs and endpoints, see APIs supported and Endpoints supported.

Endpoint Model ID In-Region endpoint URL Geo inference ID Global inference ID
bedrock-runtime zai.glm-5.3 https://bedrock-runtime.{region}.amazonaws.com us.zai.glm-5.3 global.zai.glm-5.3

For example, if region is us-east-1 (N. Virginia), then the bedrock-runtime endpoint URL will be "https://bedrock-runtime.us-east-1.amazonaws.com".

Service Tiers

Amazon Bedrock offers multiple service tiers to match your workload requirements. Standard provides pay-per-token access with no commitment (set "service_tier": "default" or omit the field). Priority delivers the fastest response times for a price premium (set "service_tier": "priority"). Flex provides lower-cost access for flexible, non-time-sensitive workloads (set "service_tier": "flex"). Reserved provides dedicated throughput with a term commitment for predictable workloads; it is set at the account level rather than per request (contact your AWS account team to enable). For more information, see service tiers.

Standard Priority Flex Reserved
Green circle with white checkmark icon. Green circle with white checkmark icon. Green circle with white checkmark icon. Red circle with white X icon indicating error, cancel, or close action.

Regional Availability

Regional availability at a glance

Amazon Bedrock offers three inference options: In-Region keeps requests within a single Region for strict compliance, Geo Cross-Region routes across Regions within a geography (such as US, EU, and APAC) while respecting data residency, and Global Cross-Region routes anywhere worldwide when there are no residency constraints. Refer to the Regional availability by models page for more details.

GLM 5.3 supports Geo Cross-Region inference from the US Regions and Global Cross-Region inference from the Regions listed in the following table. In-Region inference is not supported.

Region In-Region Geo Global
us-east-1 (N. Virginia)not-supportedsupportedsupported
us-east-2 (Ohio)not-supportedsupportedsupported
us-west-1 (N. California)not-supportedsupportedsupported
us-west-2 (Oregon)not-supportedsupportedsupported
ca-central-1 (Canada)not-supportednot-supportedsupported
ca-west-1 (Calgary)not-supportednot-supportedsupported
eu-central-1 (Frankfurt)not-supportednot-supportedsupported
eu-central-2 (Zurich)not-supportednot-supportedsupported
eu-north-1 (Stockholm)not-supportednot-supportedsupported
eu-south-1 (Milan)not-supportednot-supportedsupported
eu-south-2 (Spain)not-supportednot-supportedsupported
eu-west-1 (Ireland)not-supportednot-supportedsupported
eu-west-2 (London)not-supportednot-supportedsupported
eu-west-3 (Paris)not-supportednot-supportedsupported
ap-east-2 (Taipei)not-supportednot-supportedsupported
ap-northeast-1 (Tokyo)not-supportednot-supportedsupported
ap-northeast-2 (Seoul)not-supportednot-supportedsupported
ap-northeast-3 (Osaka)not-supportednot-supportedsupported
ap-south-1 (Mumbai)not-supportednot-supportedsupported
ap-south-2 (Hyderabad)not-supportednot-supportedsupported
ap-southeast-1 (Singapore)not-supportednot-supportedsupported
ap-southeast-2 (Sydney)not-supportednot-supportedsupported
ap-southeast-3 (Jakarta)not-supportednot-supportedsupported
ap-southeast-4 (Melbourne)not-supportednot-supportedsupported
ap-southeast-5 (Malaysia)not-supportednot-supportedsupported
ap-southeast-6 (New Zealand)not-supportednot-supportedsupported
ap-southeast-7 (Thailand)not-supportednot-supportedsupported
il-central-1 (Tel Aviv)not-supportednot-supportedsupported
af-south-1 (Cape Town)not-supportednot-supportedsupported
mx-central-1 (Mexico)not-supportednot-supportedsupported
sa-east-1 (São Paulo)not-supportednot-supportedsupported

Quotas and Limits

Your AWS account has default quotas to maintain the performance of the service and to ensure appropriate usage of Amazon Bedrock. The default quotas assigned to an account might be updated depending on regional factors, payment history, fraudulent usage, and/or approval of a quota increase request. For more information, see Quotas for Amazon Bedrock documentation and see the limits for the model.

Sample Code

Step 1 - AWS Account: If you have an AWS account already, skip this step. If you are new to AWS, sign up for an AWS account.

Step 2 - API key: Go to the Amazon Bedrock console and generate a long-term API key.

Step 3 - Get the SDK: To use this getting started guide, you must have Python already installed. Then install the relevant software depending on the APIs you are using.

Chat Completions API
pip install boto3 openai
Invoke/Converse API
pip install boto3

Step 4 - Set environment variables: Configure your environment to use the API key for authentication.

Chat Completions API
OPENAI_API_KEY="<provide your Bedrock API key>" OPENAI_BASE_URL="https://bedrock-runtime.<your-region>.amazonaws.com/openai/v1"
Invoke/Converse API
AWS_BEARER_TOKEN_BEDROCK="<provide your Bedrock API key>"

Step 5 - Run your first inference request: Save the file as bedrock-first-request.py

Chat Completions API
from openai import OpenAI client = OpenAI() response = client.chat.completions.create( model="us.zai.glm-5.3", messages=[{"role": "user", "content": "Can you explain the features of Amazon Bedrock?"}] ) print(response)
Invoke API
import json import boto3 client = boto3.client('bedrock-runtime', region_name='us-east-1') response = client.invoke_model( modelId='us.zai.glm-5.3', body=json.dumps({ 'messages': [{ 'role': 'user', 'content': 'Can you explain the features of Amazon Bedrock?'}], 'reasoning_effort': 'max', 'max_tokens': 1024 }) ) print(json.loads(response['body'].read()))
Converse API
import boto3 client = boto3.client('bedrock-runtime', region_name='us-east-1') response = client.converse( modelId='us.zai.glm-5.3', messages=[ { 'role': 'user', 'content': [{'text': 'Can you explain the features of Amazon Bedrock?'}] } ] ) print(response)