Skip to content

GLM 5.3 Flash

Tuned for throughput. The right default for chat surfaces, classification, extraction and high-volume background jobs.

Zhipu AIFastCodingavailable

Model ID

Pass this as the model field on every request.

modelglm-5.3-flash
Provider
Zhipu AI
Context window
256K tokens
Max output
64K tokens
Access
All plans
Best for
Low-latency general inference

Endpoints

Both request shapes are accepted for this model.

  • /v1/chat/completionsOpenAI chat completions format
    available
  • /v1/messagesAnthropic message format
    available
Base URLhttps://api.prbycode.com/v1

At a glance

The numbers most teams compare before switching a route.

Context256K tokens
Input$0.25 / 1M
Output$1.10 / 1M
SpeedFast
StreamingSupported

Compare with

Other models in the catalog, side by side.

ModelContextInput / 1MOutput / 1M
Claude Sonnet 51M$3.00$15.00Compare
GLM 5.2256K$0.60$2.20Compare
DeepSeek V4 Pro256K$0.90$3.40Compare
DeepSeek V4.1 Flash256K$0.20$0.80Compare
  • Claude Sonnet 51M context · $3.00 in / $15.00 out
  • GLM 5.2256K context · $0.60 in / $2.20 out
  • DeepSeek V4 Pro256K context · $0.90 in / $3.40 out
  • DeepSeek V4.1 Flash256K context · $0.20 in / $0.80 out