AWS Brings Z.ai's GLM-5.3 to Amazon Bedrock With 1M Token Context

Z.ai's 744B-parameter flagship for coding and agents lands on AWS with a 1M-token context, prompt caching, and cross-Region-only routing.

·
·
·
AWS Brings Z.ai's GLM-5.3 to Amazon Bedrock With 1M Token Context
Read4 min
TypeNews
TopicLlms · Api
  • Z.ai's GLM-5.3 is live on Amazon Bedrock for eligible enterprise customers.
  • 744B-parameter MoE with ~40B active, 1M-token context, 128K max output tokens.
  • Same base model as GLM-5.2; gains come entirely from scaled post-training.
  • Cross-Region inference only via us.zai.glm-5.3 or global.zai.glm-5.3 profiles.
  • Supports prompt caching, Guardrails, Agents, Flows, function calling, structured JSON output.
  • AWS reportedly shares per-call revenue with Z.ai under a new overseas monetization arrangement.

AWS adds GLM-5.3 to Amazon Bedrock

Amazon Web Services added Z.ai’s GLM-5.3 to Amazon Bedrock on October 5, giving eligible customers managed access to the open-weight coding and agent model through AWS APIs. The release brings a 1 million-token context window, up to 128,000 output tokens, selectable reasoning effort, prompt caching, and cross-Region inference.

Before this release, Bedrock’s Z.ai catalog had fallen behind the vendor’s model lineup. GLM-5.2 never reached the service, while six releases across three model families remained unavailable, with the oldest approaching three months. Bedrock teams can now evaluate Z.ai’s current flagship while retaining AWS billing, identity controls, and deployment tooling.

One million tokens, 40 billion active parameters

GLM-5.3 uses a Mixture-of-Experts architecture with 744 billion total parameters and about 40 billion activated for each token. The model selects a subset of specialized parameters as it processes text or code, reducing the compute required for each token compared with activating the full network.

Z.ai built GLM-5.3 on the same base model as GLM-5.2 and attributes the improvements to scaled post-training, the stages that refine tool use, instruction following, and multi-step reasoning after pretraining. The model card targets repository-scale refactoring, terminal and CLI execution, infrastructure diagnosis, and engineering tasks that span many tool calls.

The context window can hold large codebases, system prompts, and accumulated tool traces in one request. Actual capacity varies with programming language, prompt structure, generated output, and the history retained by an agent.

Routing starts with an inference profile

Bedrock serves GLM-5.3 through cross-Region inference profiles, which route requests among supported AWS Regions for capacity and availability. In-Region on-demand inference is unavailable, so applications cannot pin model execution to a single Region.

Area Details
Account access Limited to eligible customers; teams should verify access in each target AWS account.
US profile us.zai.glm-5.3
Global profile global.zai.glm-5.3, callable from supported Regions across the US, EU, Asia-Pacific, South America, and Africa.
APIs Invoke, Converse, Chat Completions, and Responses.
Output controls Function calling, structured JSON, streamed output, and streamed reasoning tokens.

The native InvokeModel API accepts the inference profile as its model ID. A minimal Python request looks like this:

haskell
import json

import boto3

client = boto3.client(
    "bedrock-runtime",
    region_name="us-east-1",
)

response = client.invoke_model(
    modelId="us.zai.glm-5.3",
    contentType="application/json",
    accept="application/json",
    body=json.dumps({
        "messages": [
            {
                "role": "user",
                "content": "Refactor this repository...",
            }
        ],
        "reasoning_effort": "max",
        "max_tokens": 1024,
    }),
)

result = json.loads(response["body"].read())
print(result)

Caching favors long agent loops

Bedrock enables prompt caching by default and accepts explicit cache checkpoints after at least 1,024 tokens, with a minimum 30-minute time to live. Agents that repeatedly send the same system prompt, repository snapshot, or tool instructions can reuse cached prefixes and reduce repeated input processing.

Supported Unavailable
Bedrock Guardrails Knowledge Bases
Bedrock Agents Intelligent prompt routing
Bedrock Flows Token counting API
Model evaluation

Long contexts and maximum reasoning effort can increase token consumption and latency, making task-level benchmarks more useful than headline context limits. AWS publishes current input, output, and caching rates on its Bedrock pricing page.

Open weights come with conditions

Z.ai distributes GLM-5.3 model files under an open-weight license, while Bedrock provides a managed endpoint and handles model serving. Applications call the service through AWS APIs rather than deploying the weights inside their own AWS accounts.

The GLM-5.3 license adds a condition for companies with more than $10 billion in revenue: they must pass a Z.ai security review before selling access to the model. The clause primarily affects hyperscalers and large inference vendors, while teams planning separate hosting or resale offerings should review the license for their use case.

Best fits and hard limits

GLM-5.3’s Bedrock deployment suits several specific workload patterns:

  • Coding agents that make many tool calls and need repository contents, instructions, and execution traces in context.
  • Refactoring and migration pipelines that benefit from adjustable reasoning effort on each request.
  • Agent loops that can reuse large cached prompt prefixes.
  • Teams moving from third-party inference to consolidate model access, billing, and governance in Bedrock.

Workloads with strict single-Region residency requirements cannot use the current cross-Region-only deployment, and applications built around Bedrock Knowledge Bases need a separate retrieval layer. Account eligibility also requires confirmation before teams commit to the model.

Production evaluations should use complete repository tasks and record tool-call success, wall-clock latency, input and output tokens, reasoning usage, and cache hits. Teams should also verify the inference profile’s permitted Regions against their security and data-residency policies.

Trending
  • No trending articles

Comments

avatar

Next Reads