Z.AI Models

71 modelsGeneral models free to startUp to 1.05M contextOfficial site

Usage

Last 30 days · 2026-08-09 to 2026-09-07

Tokens

731B

Requests

4.6M

Models in use

59 of 71

Tokens per day, stacked by model

025.8B51.6B08-0908-1608-2308-3009-062026-08-09 — 5,851,429,285 tokens glm-5.2: 3,833,910,675 42 more models: 746,759,725 coding-glm-5.2: 504,978,785 coding-glm-4.7: 498,967,700 coding-glm-5: 266,812,4002026-08-10 — 4,734,214,155 tokens glm-5.2: 2,408,502,230 42 more models: 766,239,685 coding-glm-5.2: 650,831,420 coding-glm-4.7: 459,578,135 coding-glm-5: 449,062,6852026-08-11 — 9,145,389,365 tokens coding-glm-5.2: 3,336,317,865 coding-glm-5: 2,428,312,195 glm-5.2: 1,621,252,820 42 more models: 1,302,097,535 coding-glm-4.7: 457,408,9502026-08-12 — 28,288,549,480 tokens glm-5.2: 12,373,654,255 coding-glm-5.2: 10,343,943,815 coding-glm-5: 3,735,401,105 42 more models: 1,345,942,280 coding-glm-4.7: 489,608,0252026-08-13 — 16,976,099,525 tokens coding-glm-5.2: 8,971,879,860 coding-glm-5: 3,790,038,285 glm-5.2: 2,498,606,080 42 more models: 1,107,224,900 coding-glm-4.7: 608,350,4002026-08-14 — 13,979,625,255 tokens coding-glm-5.2: 5,237,336,550 glm-5.2: 2,869,854,170 coding-glm-5: 2,754,609,075 coding-glm-5.3: 1,314,751,205 42 more models: 1,157,377,670 coding-glm-4.7: 645,696,5852026-08-15 — 8,927,698,295 tokens coding-glm-5.2: 2,894,599,215 coding-glm-5.3: 2,712,119,810 coding-glm-4.7: 971,683,360 42 more models: 966,686,285 glm-5.2: 736,207,660 coding-glm-5: 646,401,9652026-08-16 — 10,858,625,150 tokens coding-glm-5.3: 6,240,755,890 42 more models: 1,809,306,380 glm-5.2: 1,277,709,835 coding-glm-5: 968,461,655 coding-glm-4.7: 550,792,815 coding-glm-5.2: 11,598,5752026-08-17 — 10,885,006,880 tokens coding-glm-5.3: 4,856,441,140 glm-5.2: 3,542,775,825 42 more models: 989,148,555 coding-glm-5: 846,574,650 coding-glm-4.7: 534,609,745 coding-glm-5.2: 109,125,935 glm-5.3: 6,331,0302026-08-18 — 12,680,941,755 tokens coding-glm-5.3: 5,588,652,325 glm-5.2: 3,295,448,645 42 more models: 1,064,898,990 coding-glm-4.7: 801,399,540 coding-glm-5: 755,660,520 glm-5.3: 615,618,425 coding-glm-5.2: 559,263,3102026-08-19 — 16,957,620,515 tokens coding-glm-5.3: 10,813,449,925 glm-5.3: 2,034,350,875 glm-5.2: 1,311,421,615 42 more models: 1,183,430,830 coding-glm-4.7: 774,367,860 coding-glm-5: 669,518,985 coding-glm-5.2: 171,080,4252026-08-20 — 18,954,357,385 tokens coding-glm-5.3: 10,854,316,955 glm-5.3: 3,115,697,245 glm-5.2: 1,984,767,335 42 more models: 1,586,142,390 coding-glm-5: 700,718,890 coding-glm-4.7: 601,639,450 coding-glm-5.2: 111,075,1202026-08-21 — 14,319,820,885 tokens coding-glm-5.3: 6,400,303,505 glm-5.3: 3,551,356,595 42 more models: 1,852,746,185 glm-5.2: 914,098,525 coding-glm-4.7: 836,180,500 coding-glm-5: 752,224,190 coding-glm-5.2: 12,911,3852026-08-22 — 11,655,525,000 tokens glm-5.3: 5,314,556,755 coding-glm-5.3: 3,263,978,225 42 more models: 1,345,746,230 coding-glm-5: 749,770,845 coding-glm-4.7: 493,553,955 glm-5.2: 432,391,315 coding-glm-5.2: 55,527,6752026-08-23 — 7,972,566,555 tokens glm-5.3: 5,254,336,005 42 more models: 1,288,402,255 coding-glm-5: 729,098,185 coding-glm-4.7: 437,156,505 glm-5.2: 209,659,595 coding-glm-5.2: 53,914,0102026-08-24 — 12,228,887,180 tokens glm-5.3: 7,151,130,030 42 more models: 1,539,399,225 coding-glm-5.3: 1,057,873,455 coding-glm-5.2: 854,044,745 glm-5.2: 677,333,920 coding-glm-5: 638,361,775 coding-glm-4.7: 310,744,0302026-08-25 — 13,624,077,940 tokens glm-5.3: 5,884,650,075 coding-glm-5.3: 2,837,134,225 42 more models: 1,835,901,105 glm-5.2: 1,268,252,005 coding-glm-5: 922,107,850 coding-glm-4.7: 503,760,210 coding-glm-5.2: 372,272,4702026-08-26 — 13,319,822,160 tokens glm-5.3: 6,189,403,720 coding-glm-5.3: 2,201,414,415 glm-5.2: 1,875,141,405 42 more models: 1,727,162,710 coding-glm-4.7: 574,062,960 coding-glm-5: 564,909,435 coding-glm-5.2: 180,769,625 glm-5.3-flash: 6,957,8902026-08-27 — 42,417,976,010 tokens glm-5.3-flash: 18,544,831,100 coding-glm-5.3: 8,942,687,205 42 more models: 6,400,467,310 glm-5.3: 5,401,025,930 coding-glm-5.3-flash: 1,270,865,940 coding-glm-5: 675,621,425 coding-glm-4.7: 581,479,780 glm-5.2: 579,913,790 coding-glm-5.2: 21,083,5302026-08-28 — 49,882,099,935 tokens glm-5.3-flash: 24,967,167,410 glm-5.3: 6,644,999,460 coding-glm-5.3: 6,473,492,100 42 more models: 5,598,689,855 coding-glm-5.3-flash: 4,668,933,520 coding-glm-4.7: 702,247,515 coding-glm-5: 638,659,035 glm-5.2: 146,821,020 coding-glm-5.2: 41,090,0202026-08-29 — 42,255,269,785 tokens glm-5.3-flash: 19,797,770,380 coding-glm-5.3: 9,379,725,515 coding-glm-5.3-flash: 4,368,748,005 glm-5.3: 3,828,966,350 42 more models: 3,664,758,820 coding-glm-5: 546,139,410 coding-glm-4.7: 468,092,615 glm-5.2: 165,515,320 coding-glm-5.2: 35,553,3702026-08-30 — 38,240,889,900 tokens glm-5.3-flash: 16,887,788,695 coding-glm-5.3-flash: 7,075,907,190 coding-glm-5.3: 6,985,081,055 42 more models: 3,037,626,815 glm-5.3: 2,517,170,120 coding-glm-4.7: 657,528,045 coding-glm-5: 509,579,135 coding-glm-5.2: 329,217,110 glm-5.2: 240,991,7352026-08-31 — 43,047,594,125 tokens glm-5.3-flash: 30,310,025,780 glm-5.3: 6,139,648,995 coding-glm-5.3: 2,144,963,070 42 more models: 1,825,336,025 coding-glm-5.3-flash: 1,676,527,185 coding-glm-5: 393,732,375 coding-glm-4.7: 304,180,545 glm-5.2: 247,680,685 coding-glm-5.2: 5,499,4652026-09-01 — 51,461,371,390 tokens glm-5.3-flash: 31,654,991,910 glm-5.3: 9,091,797,005 coding-glm-5.3-flash: 4,915,068,365 42 more models: 2,373,256,310 coding-glm-5.3: 1,835,721,165 glm-5.2: 607,777,805 coding-glm-4.7: 517,627,840 coding-glm-5: 451,663,420 coding-glm-5.2: 13,467,5702026-09-02 — 46,666,343,135 tokens glm-5.3-flash: 26,759,614,510 coding-glm-5.3-flash: 7,740,056,995 glm-5.3: 5,170,680,560 coding-glm-5.3: 3,393,608,975 42 more models: 2,079,838,445 coding-glm-4.7: 858,361,495 coding-glm-5: 555,869,820 glm-5.2: 101,757,240 coding-glm-5.2: 6,555,0952026-09-03 — 51,641,943,395 tokens glm-5.3-flash: 25,230,128,915 coding-glm-5.3-flash: 12,343,929,790 coding-glm-5.3: 6,869,311,775 glm-5.3: 3,979,112,850 42 more models: 1,853,474,590 coding-glm-4.7: 649,069,910 coding-glm-5: 479,669,045 glm-5.2: 218,232,175 coding-glm-5.2: 19,014,3452026-09-04 — 46,371,361,975 tokens glm-5.3-flash: 19,662,008,285 coding-glm-5.3: 9,763,910,860 glm-5.3: 9,293,907,790 coding-glm-5.3-flash: 4,901,606,785 42 more models: 1,549,515,290 coding-glm-4.7: 602,820,105 coding-glm-5: 414,301,745 glm-5.2: 117,095,820 coding-glm-5.2: 66,195,2952026-09-05 — 29,861,185,965 tokens glm-5.3-flash: 21,927,219,175 glm-5.3: 3,128,027,890 coding-glm-5.3-flash: 1,852,618,190 42 more models: 1,575,108,090 coding-glm-5.3: 685,183,950 coding-glm-4.7: 269,966,815 coding-glm-5: 234,625,675 glm-5.2: 151,125,030 coding-glm-5.2: 37,311,1502026-09-06 — 23,964,835,395 tokens glm-5.3-flash: 15,843,079,290 coding-glm-5.3-flash: 2,664,937,935 glm-5.3: 2,492,985,755 42 more models: 1,417,542,445 coding-glm-5.3: 1,056,058,050 glm-5.2: 253,601,235 coding-glm-5: 216,045,370 coding-glm-5.2: 20,170,455 coding-glm-4.7: 414,8602026-09-07 — 34,215,316,590 tokens glm-5.3-flash: 19,059,816,875 glm-5.3: 8,506,221,890 coding-glm-5.3-flash: 3,673,599,520 coding-glm-5.3: 1,706,993,930 42 more models: 909,134,640 glm-5.2: 170,388,190 coding-glm-5: 167,196,945 coding-glm-5.2: 21,461,830 coding-glm-4.7: 502,770
  • glm-5.3-flash
  • coding-glm-5.3
  • glm-5.3
  • coding-glm-5.3-flash
  • glm-5.2
  • coding-glm-5.2
  • coding-glm-5
  • coding-glm-4.7
  • 42 more models

Which models that traffic went to

  1. GLM 5.3 Flash37.0%271B
  2. Coding GLM 5.316.0%117B
  3. GLM 5.314.4%105B
  4. Coding GLM 5.3 Flash7.8%57.2B
  5. GLM 5.26.3%46.1B
  6. Coding GLM 5.24.8%35B
  7. Coding GLM 53.8%27.7B
  8. Coding GLM 4.72.2%16.2B
  9. 42 more models7.6%55.9B

Share of 731B tokens. 9 models with traffic report no token counts and cannot be ranked here, including coding-glm-5-turbo-free and glm-4-flash — they are in the request view.

The two views disagree on purpose: a model can take a large share of the calls and a small share of the tokens — many short requests — or the reverse. Which one matters depends on whether your cost is driven by call volume or by prompt length. Measured on AIHubMix over the last 30 days, counting the 71 model IDs listed on this page; traffic routed through upstream-specific IDs that are not in the public catalog is not included.

All 71 Z.AI Models

Open in model list
Z.AI models on AIHubMix with input and output modalities, context length, maximum output, price per million tokens including cache read and cache write rates, and measured throughput and latency.
Modalities
coding-glm-5.3-freeTakes , returns text.1.05MFreeFree/M
ox-alphaTakes text, vision, returns text.1.05MFreeFree/M
coding-glm-5.3Takes text, returns text.1.05M$0.06$0.22/M$0.015/M
glm-5.3-flashTakes text, vision, video, returns text.1.05M128K$0.1127$0.3944/M$0.0282/M90 tok/s2.76 s
glm-5.3Takes text, returns text.1.05M128K$1.1268$3.9438/M$0.2817/M29 tok/s2.64 s
coding-glm-5.2-freeTakes text, returns text.1MFreeFree/M
coding-glm-5.3-flash-free1MFreeFree/M
coding-glm-5.3-flash1M128K$0.0282$0.0986/M$0.007/M
coding-glm-5.2Takes text, returns text.1M$0.06$0.22/M
glm-5.2Takes text, returns text.1M128K$1.1268$3.9438/M$0.2817/M36 tok/s0.64 s
glm-5.2-fast-previewTakes text, returns text.1M128K$2.254$7.889/M$0.5635/M40 tok/s1.60 s
coding-glm-5-turbo-freeTakes text, returns text.205KFreeFree/M
glm-4.6Takes text, returns text.205K131KFreeFree/MFree/M40 tok/s1.71 s
glm-5-turboTakes text, returns text.205K$1.2$3.9996/M$0.24/M19 tok/s6.02 s
coding-glm-4.6-freeTakes text, returns text.200K128KFreeFree/M
coding-glm-5-freeTakes text, returns text.200KFreeFree/M
coding-glm-5.1-freeTakes text, returns text.200KFreeFree/M
glm-5Takes text, returns text.200KFreeFree/MFree/M64 tok/s0.88 s
cc-glm-5.1Takes text, returns text.200K$0.06$0.22/M
coding-glm-5.1Takes text, returns text.200K$0.06$0.22/M
glm-4.7Takes text, returns text.200K128K$0.274$1.0959/M$0.0548/M27 tok/s12.96 s
glm-5v-turboTakes text, vision, video, returns text.200K128K$0.7042$3.0985/M$0.169/M35 tok/s4.43 s
glm-5.1Takes text, returns text.200K128K$0.845$3.38/M$0.1831/M22 tok/s1.27 s
coding-glm-4.5-airTakes text. Output modality not published.131K$0.014$0.084/M
glm-4.5-airTakes text. Output modality not published.131K98K$0.14$0.84/M63 tok/s0.73 s
glm-4.5Takes text. Output modality not published.131K98K$0.4$1.6/M107 tok/s0.46 s
glm-4.6vTakes text, vision, video, returns text.128K$0.137$0.411/M$0.0274/M18 tok/s1.72 s
glm-4.5vTakes text, vision, video, returns text.64K16K$0.274$0.822/M81 tok/s8.45 s
glm-ocrTakes vision, returns text.32K$0.0282$0.0282/M
embedding-2Takes text. Output modality not published.8K$0.0686$0.0686/M
embedding-3Takes text. Output modality not published.8K$0.0686$0.0686/M
coding-glm-4.7-freeTakes text, returns text.FreeFree/M
glm-4.7-flash-freeTakes text, returns text.FreeFree/M
glm-imageTakes text, returns vision.FreeFree/M
Pro/THUDM/GLM-4.1V-9B-Thinking$0.04$0.16/M
THUDM/GLM-4-9B-0414$0.05$0.05/M
THUDM/GLM-Z1-9B-0414$0.05$0.05/M
cc-glm-4.6$0.06$0.22/M
cc-glm-4.7$0.06$0.22/M
cc-glm-5Takes text, returns text.$0.06$0.22/M
cc-glm-5-turboTakes text, returns text.$0.06$0.22/M
coding-glm-4.6Takes text, returns text.$0.06$0.22/M$0.011/M
coding-glm-4.7Takes text, returns text.$0.06$0.22/M$0.011/M
coding-glm-5Takes text, returns text.$0.06$0.22/M
coding-glm-5-turboTakes text, returns text.$0.06$0.22/M
THUDM/GLM-4-32B-0414$0.08$0.08/M
THUDM/GLM-Z1-32B-0414$0.08$0.08/M
glm-4-flash$0.1$0.1/M
THUDM/GLM-4.1V-9B-Thinking$0.1$0.1/M
doubao-1-5-pro-32k-250115$0.108$0.27/M
chatglm_lite$0.2858$0.2858/M
alicloud-glm-4.7$0.411$1.9178/M$0.411/M
alicloud-glm-5$0.5634$2.5353/M$0.1127/M
doubao-1-5-pro-256k-250115$0.684$1.2312/M
glm-3-turbo$0.71$0.71/M
chatglm_std$0.7144$0.7144/M
chatglm_turbo$0.7144$0.7144/M
glm-4.5-airxTakes text. Output modality not published.$1.1$4.51/M$0.22/M
zai-glm-5-turboTakes , returns text.$1.2$3.9996/M$0.24/M
cloudflare-glm-5.2Takes , returns text.$1.4$4.4002/M$0.2604/M
chatglm_pro$1.4286$1.4286/M
glm-4v-plus$2$2/M
glm-zero-preview$2$2/M
glm-4.5-xTakes text. Output modality not published.$2.2$8.91/M$0.44/M1 tok/s0.59 s
cbs-glm-4.7$2.25$2.75/M
glm-4-plus$8$8/M
cogview-3-plus$10$10/M
glm-4$14.2$14.2/M
glm-4v$14.2$14.2/M
code-davinci-edit-001$20$20/M
cogview-3$35.5$35.5/M

Prices are USD per million tokens; cache read and cache write are the rates for prompt-cache hits and for writing a prompt into the cache. Throughput and latency are measured on AIHubMix — the same figures the model detail page shows — not vendor claims. A dash means the catalog does not publish that field for that model, which is not the same as the model not supporting it.

Z.AI on AIHubMix

Which Z.AI model should I start with?

coding-glm-4.6-free is free on input — the cheapest entry here that declares tool calling, and it carries a 200K context. Move up to cogview-3 when answer quality matters more than cost, or to coding-glm-5.3-free for long-form reasoning.

Which of these models reason before answering?

30 of the 71 models here declare a reasoning phase — they work through the problem before producing an answer, which helps on multi-step problems at the cost of extra output tokens. Use the Reasoning filter above the table to see them. The catalog does not record anything further about how they differ, so this page does not sort them into families.

Why are there several entries for the same model?

Because each row is a route you can call, not a model release. Some IDs name an upstream (azure-, alicloud-, cc-), some are the open-weight repository form (THUDM/…), and some differ only in capitalisation, kept so older integrations keep working.

The catalog does not carry a field saying which of those a given row is, so this page does not sort them into buckets it would have to invent. Every row shows that route’s own price, context and speed — compare those directly, and open a model to see the upstreams that serve it.

How is cached input billed?

The Cache read column is the rate for input tokens served from the prompt cache — for example coding-glm-4.6 bills cache hits at 18.33% of the input rate and coding-glm-4.7 bills cache hits at 18.33% of the input rate. Cache write is the surcharge for putting a prompt into the cache in the first place, and only a few upstreams bill it separately. A dash in either column means the catalog carries no cache rate for that model, so plan on paying the full input rate.

Do I need a separate Z.AI account?

No. One AIHubMix key covers every model on this page, and switching between them is a change to the model string — billing, rate limits, and logs stay in one place.

Start calling Z.AI in one line

One key, one endpoint, 880 models across 38 model authors.