Error 403: what to do when glm-5.1 fails

A 403 means you are authenticated but not allowed: the key is valid, yet it may not call this model or this group.

403 Forbidden means the identity was recognised but access is denied. The difference from 401: 401 asks who you are, 403 says you lack permission. Usually the model needs to be added to the allowed list or group bound to your key.

glm-5.1 is served by Z.AI. Everything on this page — triggers, fixes and measured data — is compiled from the real runtime behaviour of this model at the gateway layer.

At this gateway, the most common trigger is: This key is not authorised for this model. The recommended first action is: Add this model to the key allowlist on the Tokens page.

Common causes

  • This key is not authorised for this model
  • The model is not in the group the key is bound to
  • The key has an IP allowlist that excludes your current IP
  • The model was retired or requires higher permissions

How to fix

  • Add this model to the key allowlist on the Tokens page
  • Confirm the key group includes this model
  • Review the IP allowlist settings
  • Switch to a model you are allowed to call

Retry with exponential backoff

The snippet below retries when glm-5.1 returns 403, up to 5 attempts, with an increasing wait plus random jitter so concurrent calls do not retry in lockstep. Read the base URL and API key from environment variables — never hardcode them.

import os, time, random
import requests

BASE  = os.getenv("OPENAI_BASE_URL")   # e.g. https://<your-gateway>/v1
KEY   = os.getenv("OPENAI_API_KEY")
MODEL = 'glm-5.1'


def chat(messages, retries=5):
    """Retry with exponential backoff + jitter."""
    for i in range(retries):
        try:
            r = requests.post(
                BASE + "/chat/completions",
                headers={"Authorization": "Bearer " + KEY},
                json={"model": MODEL, "messages": messages, "stream": True},
                timeout=60,
            )
            if r.status_code == 429 or r.status_code >= 500:
                time.sleep(min(2 ** i + random.uniform(0, 1), 30))
                continue
            r.raise_for_status()
            return r.json()
        except requests.exceptions.Timeout:
            time.sleep(min(2 ** i + random.uniform(0, 1), 30))
    raise RuntimeError("gave up after " + str(retries) + " retries")


print(chat([{"role": "user", "content": "hello"}]))

Key facts for this model

API endpointhttps://api.airai.cc/v1
OpenAI-compatibleOpenAI-compatible
VendorZ.AI
Context200K
CapabilitiesReasoning, Tools, Open Weights
API formatsopenai, openai-response, openai-response-compact, anthropic, gemini, openai-alpha-search
Billing formulap * 1.4 + cr * 0.26 + cc * 0 + c * 4.4

FAQ

Will I be charged when glm-5.1 returns 403?

No charge — only output actually produced counts toward usage. Input price is about $1.40 per million tokens. Check your balance and rate limits in the console before debugging code.

Does sharing one key across several services make 403 more likely?

It is mainly a quota matter, not a fault in the model itself. Billing follows p * 1.4 + cr * 0.26 + cc * 0 + c * 4.4, so no output means no charge. With billing p * 1.4 + cr * 0.26 + cc * 0 + c * 4.4, the cost of long output comes mostly from output tokens. Check your balance and rate limits in the console before debugging code. When estimating cost from p * 1.4 + cr * 0.26 + cc * 0 + c * 4.4, include the retry budget.

Can async or batch processing avoid 403?

Concurrency and timeouts are the real variables here, not the model itself. This model has a 200K context window and comes from Z.AI. A 200K context means long inputs add noticeably to first-token latency. For long outputs, raise the timeout to 60 seconds or more. Use batching or a queue to smooth peaks — steadier than raising concurrency on the fly.

403 keeps recurring on glm-5.1 — how do I tell whether it is the model or my account?

This error is unrelated to model capability; it is a gateway-layer issue. The capability tags for this model are Reasoning, Tools, Open Weights. If it persists for several minutes, contact platform support to confirm upstream status. Decide account-level versus model-level first; the two need completely different handling.

Other errors on this model

Other models with the same error

Data updated: 2026-10-10 15:45

Technical SupportLive Support
Back to Top