Z.AI Models

71 modelsGeneral models free to startUp to 1.05M contextOfficial site ↗

Usage

Last 30 days · 2026-09-03 to 2026-10-02

Tokens

902B

Requests

6.4M

Models in use

59 of 71

Tokens per day, stacked by model

025.8B51.6B09-0309-1009-1709-2410-012026-09-03 — 51,641,943,190 tokens glm-5.3-flash: 25,230,128,915 coding-glm-5.3-flash: 12,343,929,790 coding-glm-5.3: 6,869,311,775 glm-5.3: 3,979,112,850 39 more models: 1,945,448,050 coding-glm-5.3-free: 747,105,115 coding-glm-5.3-flash-free: 308,674,520 glm-5.2: 218,232,1752026-09-04 — 46,370,617,930 tokens glm-5.3-flash: 19,662,008,285 coding-glm-5.3: 9,763,910,860 glm-5.3: 9,293,907,790 coding-glm-5.3-flash: 4,901,606,785 39 more models: 1,728,253,770 coding-glm-5.3-free: 681,732,660 coding-glm-5.3-flash-free: 222,101,960 glm-5.2: 117,095,8202026-09-05 — 29,861,168,910 tokens glm-5.3-flash: 21,927,219,175 glm-5.3: 3,128,027,890 coding-glm-5.3-flash: 1,852,618,190 39 more models: 1,590,800,550 coding-glm-5.3: 685,183,950 coding-glm-5.3-flash-free: 287,493,630 coding-glm-5.3-free: 238,700,495 glm-5.2: 151,125,0302026-09-06 — 23,964,673,455 tokens glm-5.3-flash: 15,843,079,290 coding-glm-5.3-flash: 2,664,937,935 glm-5.3: 2,492,985,755 39 more models: 1,183,759,895 coding-glm-5.3: 1,056,058,050 coding-glm-5.3-free: 296,578,685 glm-5.2: 253,601,235 coding-glm-5.3-flash-free: 173,672,6102026-09-07 — 34,211,324,710 tokens glm-5.3-flash: 19,059,816,875 glm-5.3: 8,506,221,890 coding-glm-5.3-flash: 3,673,599,520 coding-glm-5.3: 1,706,993,930 39 more models: 710,335,795 coding-glm-5.3-free: 203,013,875 coding-glm-5.3-flash-free: 180,954,635 glm-5.2: 170,388,1902026-09-08 — 40,512,087,435 tokens glm-5.3-flash: 31,111,412,135 glm-5.3: 4,765,267,430 coding-glm-5.3-flash: 1,836,691,175 coding-glm-5.3: 1,186,262,485 39 more models: 759,650,485 coding-glm-5.3-free: 489,010,705 coding-glm-5.3-flash-free: 282,641,135 glm-5.2: 81,151,8852026-09-09 — 37,302,208,485 tokens glm-5.3-flash: 28,819,987,670 glm-5.3: 2,668,476,660 coding-glm-5.3-flash: 2,630,894,580 coding-glm-5.3: 1,787,867,080 coding-glm-5.3-free: 665,144,695 39 more models: 464,953,010 coding-glm-5.3-flash-free: 231,594,160 glm-5.2: 33,290,6302026-09-10 — 36,761,777,540 tokens glm-5.3-flash: 20,838,310,355 coding-glm-5.3: 6,256,329,915 coding-glm-5.3-flash: 4,554,588,335 glm-5.3: 3,918,748,480 39 more models: 509,657,875 coding-glm-5.3-free: 337,835,180 glm-5.2: 179,967,880 coding-glm-5.3-flash-free: 166,339,5202026-09-11 — 27,464,306,905 tokens glm-5.3-flash: 13,499,143,900 coding-glm-5.3: 6,302,106,195 coding-glm-5.3-flash: 4,176,616,120 glm-5.3: 2,423,067,985 coding-glm-5.3-free: 387,048,340 39 more models: 296,418,520 glm-5.2: 195,702,255 coding-glm-5.3-flash-free: 184,203,5902026-09-12 — 28,040,164,260 tokens glm-5.3-flash: 12,697,067,645 coding-glm-5.3: 5,913,778,670 coding-glm-5.3-flash: 5,901,329,840 glm-5.3: 2,562,207,840 39 more models: 359,770,940 coding-glm-5.3-free: 349,691,025 coding-glm-5.3-flash-free: 201,259,125 glm-5.2: 55,059,1752026-09-13 — 28,973,843,095 tokens glm-5.3-flash: 15,594,889,085 coding-glm-5.3-flash: 7,183,268,420 coding-glm-5.3: 3,550,681,185 glm-5.3: 1,597,157,395 39 more models: 482,251,180 coding-glm-5.3-free: 270,211,240 coding-glm-5.3-flash-free: 202,108,260 glm-5.2: 93,276,3302026-09-14 — 29,853,050,580 tokens glm-5.3-flash: 13,233,719,365 coding-glm-5.3-flash: 7,729,760,530 coding-glm-5.3: 3,938,101,995 glm-5.3: 3,303,513,370 39 more models: 600,610,530 glm-5.2: 428,183,875 coding-glm-5.3-free: 362,332,145 coding-glm-5.3-flash-free: 256,828,7702026-09-15 — 31,273,410,215 tokens glm-5.3-flash: 13,130,306,175 coding-glm-5.3-flash: 8,211,169,435 glm-5.3: 5,153,701,615 coding-glm-5.3: 3,269,071,695 39 more models: 456,764,150 coding-glm-5.3-free: 448,137,335 glm-5.2: 336,176,875 coding-glm-5.3-flash-free: 268,082,9352026-09-16 — 32,518,141,520 tokens glm-5.3-flash: 14,586,531,655 glm-5.3: 9,593,025,950 coding-glm-5.3-flash: 3,440,419,375 coding-glm-5.3: 3,124,363,945 39 more models: 876,271,915 coding-glm-5.3-free: 515,564,400 coding-glm-5.3-flash-free: 222,646,355 glm-5.2: 159,317,9252026-09-17 — 44,993,223,120 tokens glm-5.3-flash: 15,877,690,165 coding-glm-5.3: 11,582,806,730 glm-5.3: 10,737,226,180 coding-glm-5.3-flash: 5,313,812,065 39 more models: 556,169,210 coding-glm-5.3-free: 398,743,550 glm-5.2: 344,222,045 coding-glm-5.3-flash-free: 182,553,1752026-09-18 — 36,122,929,035 tokens coding-glm-5.3: 13,823,751,070 glm-5.3-flash: 11,001,676,670 glm-5.3: 6,410,825,905 coding-glm-5.3-flash: 3,477,628,505 39 more models: 632,507,640 coding-glm-5.3-free: 396,123,935 glm-5.2: 213,453,700 coding-glm-5.3-flash-free: 161,402,855 glm-5.3-flashx: 5,558,7552026-09-19 — 28,533,962,090 tokens coding-glm-5.3: 11,623,725,605 glm-5.3-flash: 6,963,528,085 coding-glm-5.3-flash: 4,645,912,290 glm-5.3: 2,312,909,370 glm-5.3-flashx: 1,395,826,035 coding-glm-5.3-free: 680,200,690 39 more models: 596,779,200 coding-glm-5.3-flash-free: 169,289,510 glm-5.2: 145,791,3052026-09-20 — 19,400,823,010 tokens glm-5.3-flash: 8,460,224,920 glm-5.3: 4,200,871,800 coding-glm-5.3: 2,918,865,425 glm-5.3-flashx: 2,808,218,655 39 more models: 425,424,530 glm-5.2: 277,743,675 coding-glm-5.3-free: 193,201,335 coding-glm-5.3-flash: 99,320,465 coding-glm-5.3-flash-free: 16,952,2052026-09-21 — 37,804,096,995 tokens glm-5.3-flash: 19,942,791,035 glm-5.3: 13,101,303,460 glm-5.3-flashx: 2,069,376,710 coding-glm-5.3: 1,670,930,525 glm-5.2: 430,262,135 39 more models: 385,893,010 coding-glm-5.3-flash: 144,302,820 coding-glm-5.3-free: 57,646,375 coding-glm-5.3-flash-free: 1,590,9252026-09-22 — 35,869,308,860 tokens glm-5.3: 10,710,281,705 glm-5.3-flash: 10,339,650,560 glm-5.3-flashx: 5,418,770,705 coding-glm-5.3-flash: 4,192,272,935 coding-glm-5.3: 3,633,681,365 39 more models: 719,905,980 glm-5.2: 498,997,350 coding-glm-5.3-free: 302,025,405 coding-glm-5.3-flash-free: 53,722,8552026-09-23 — 30,465,635,140 tokens glm-5.3-flash: 8,747,566,540 glm-5.3: 7,301,127,995 coding-glm-5.3: 5,410,572,435 glm-5.3-flashx: 4,369,220,515 coding-glm-5.3-flash: 3,066,424,605 glm-5.2: 745,096,625 39 more models: 485,236,315 coding-glm-5.3-free: 304,845,810 coding-glm-5.3-flash-free: 35,544,3002026-09-24 — 34,164,094,985 tokens glm-5.3-flash: 10,569,837,200 glm-5.3: 8,364,816,115 coding-glm-5.3: 5,898,677,820 coding-glm-5.3-flash: 4,785,211,620 glm-5.3-flashx: 3,327,104,625 39 more models: 736,897,350 coding-glm-5.3-free: 272,651,730 glm-5.2: 143,812,845 coding-glm-5.3-flash-free: 65,085,6802026-09-25 — 15,343,110,980 tokens glm-5.3-flash: 4,985,971,445 coding-glm-5.3-flash: 4,051,147,350 coding-glm-5.3: 2,883,232,820 glm-5.3: 967,827,430 glm-5.3-flashx: 840,149,940 glm-5.2: 666,315,440 39 more models: 595,640,090 coding-glm-5.3-free: 312,695,275 coding-glm-5.3-flash-free: 40,131,1902026-09-26 — 20,651,087,210 tokens coding-glm-5.3: 8,605,799,410 glm-5.3-flash: 6,829,929,810 coding-glm-5.3-flash: 3,320,270,070 glm-5.3: 946,614,890 39 more models: 524,390,465 coding-glm-5.3-free: 324,836,260 glm-5.2: 34,994,665 glm-5.3-flashx: 32,757,355 coding-glm-5.3-flash-free: 31,494,2852026-09-27 — 13,615,805,890 tokens glm-5.3-flash: 4,872,180,945 coding-glm-5.3: 4,716,002,910 coding-glm-5.3-flash: 2,075,529,115 glm-5.3: 1,004,714,310 39 more models: 406,373,230 coding-glm-5.3-free: 270,785,990 glm-5.3-flashx: 132,078,085 glm-5.2: 114,022,645 coding-glm-5.3-flash-free: 24,118,6602026-09-28 — 19,556,045,830 tokens coding-glm-5.3: 6,818,499,240 glm-5.3-flash: 5,431,874,900 coding-glm-5.3-flash: 3,167,034,665 glm-5.3: 2,905,327,720 39 more models: 627,319,670 coding-glm-5.3-free: 343,690,100 glm-5.3-flashx: 219,024,280 glm-5.2: 24,568,750 coding-glm-5.3-flash-free: 18,706,5052026-09-29 — 25,181,650,980 tokens coding-glm-5.3-flash: 7,876,009,835 coding-glm-5.3: 7,106,682,645 glm-5.3-flash: 5,300,363,395 glm-5.3: 2,705,659,930 glm-5.3-flashx: 1,186,217,030 39 more models: 561,614,410 coding-glm-5.3-free: 284,678,940 glm-5.2: 119,985,360 coding-glm-5.3-flash-free: 40,439,4352026-09-30 — 17,091,670,370 tokens glm-5.3-flash: 5,004,491,970 coding-glm-5.3-flash: 4,938,282,050 coding-glm-5.3: 4,289,874,190 39 more models: 997,539,985 glm-5.3-flashx: 795,160,420 glm-5.3: 637,618,345 coding-glm-5.3-free: 271,319,615 glm-5.2: 118,861,175 coding-glm-5.3-flash-free: 38,522,6202026-10-01 — 16,228,282,305 tokens glm-5.3-flash: 5,218,315,370 coding-glm-5.3-flash: 5,092,997,610 glm-5.3: 2,673,209,540 coding-glm-5.3: 1,347,471,725 39 more models: 1,161,923,770 glm-5.3-flashx: 402,956,680 coding-glm-5.3-free: 186,515,730 glm-5.2: 117,118,950 coding-glm-5.3-flash-free: 27,772,9302026-10-02 — 28,131,681,730 tokens coding-glm-5.3-flash: 16,045,784,060 coding-glm-5.3: 3,919,846,055 glm-5.3-flash: 3,697,162,225 glm-5.3: 2,986,531,205 39 more models: 665,622,040 glm-5.3-flashx: 350,196,290 coding-glm-5.3-free: 310,576,505 glm-5.2: 134,886,820 coding-glm-5.3-flash-free: 21,076,530
  • glm-5.3-flash
  • coding-glm-5.3
  • coding-glm-5.3-flash
  • glm-5.3
  • glm-5.3-flashx
  • coding-glm-5.3-free
  • glm-5.2
  • coding-glm-5.3-flash-free
  • 39 more models

Which models that traffic went to

  1. GLM 5.3 Flash44.2%398B
  2. Coding GLM 5.316.8%152B
  3. Coding GLM 5.3 Flash15.9%143B
  4. GLM 5.315.7%141B
  5. GLM 5.3 Flashx2.6%23.4B
  6. Coding GLM 5.3 (free)1.2%10.9B
  7. GLM 5.20.7%6.6B
  8. Coding GLM 5.3 Flash (free)0.5%4.1B
  9. 39 more models2.4%22B

Share of 902B tokens. 12 models with traffic report no token counts and cannot be ranked here, including coding-glm-5-turbo-free and ox-alpha — they are in the request view.

The two views disagree on purpose: a model can take a large share of the calls and a small share of the tokens — many short requests — or the reverse. Which one matters depends on whether your cost is driven by call volume or by prompt length. Measured on AIHubMix over the last 30 days, counting the 71 model IDs listed on this page; traffic routed through upstream-specific IDs that are not in the public catalog is not included.

All 71 Z.AI Models

Open in model list
Z.AI models on AIHubMix with input and output modalities, context length, maximum output, price per million tokens including cache read and cache write rates, and measured throughput and latency.
Modalities
coding-glm-5.3-freeTakes text, returns text.1.05M131KFreeFree/M—42 tok/s5.03 s
ox-alphaTakes text, vision, video, returns text.1.05M131KFreeFree/M———
coding-glm-5.3Takes text, returns text.1.05M131K$0.06$0.22/M$0.015/M40 tok/s5.11 s
glm-5.3-flashTakes text, vision, video, returns text.1.05M131K$0.1127$0.3944/M$0.0282/M369 tok/s2.78 s
glm-5.3Takes text, returns text.1.05M131K$1.1268$3.9438/M$0.2817/M49 tok/s0.59 s
coding-glm-5.2-freeTakes text, returns text.1M131KFreeFree/M—44 tok/s5.13 s
coding-glm-5.3-flash-freeTakes text, vision, video. Output modality not published.1M131KFreeFree/M—32 tok/s4.37 s
coding-glm-5.3-flashTakes text, vision, video. Output modality not published.1M131K$0.0282$0.0986/M$0.007/M28 tok/s4.09 s
coding-glm-5.2Takes text, returns text.1M131K$0.06$0.22/M—38 tok/s4.16 s
glm-5.3-flashxTakes text, vision, video, returns text.1M—$0.37$1.25/M$0.075/M363 tok/s3.31 s
glm-5.2Takes text, returns text.1M131K$1.1268$3.9438/M$0.2817/M30 tok/s0.72 s
cloudflare-glm-5.2Takes text, returns text.1M131K$1.4$4.4002/M$0.2604/M——
glm-5.2-fast-previewTakes text, returns text.1M131K$2.254$7.889/M$0.5635/M38 tok/s2.00 s
coding-glm-5-turbo-freeTakes text, returns text.205K131KFreeFree/M———
cc-glm-5-turboTakes text, returns text.205K131K$0.06$0.22/M—8 tok/s1.86 s
coding-glm-5-turboTakes text, returns text.205K131K$0.06$0.22/M———
glm-5-turboTakes text, returns text.205K131K$1.2$3.9996/M$0.24/M29 tok/s3.50 s
zai-glm-5-turboTakes text, returns text.205K131K$1.2$3.9996/M$0.24/M29 tok/s3.50 s
coding-glm-4.6-freeTakes text, returns text.200K131KFreeFree/M—28 tok/s3.16 s
coding-glm-4.7-freeTakes text, returns text.200K131KFreeFree/M—32 tok/s3.23 s
coding-glm-5-freeTakes text, returns text.200K131KFreeFree/M—30 tok/s4.88 s
coding-glm-5.1-freeTakes text, returns text.200K131KFreeFree/M—44 tok/s4.10 s
glm-4.6Takes text, returns text.200K131KFreeFree/MFree/M29 tok/s4.65 s
glm-4.7-flash-freeTakes text, returns text.200K131KFreeFree/M—37 tok/s22.43 s
cc-glm-5Takes text, returns text.200K131K$0.06$0.22/M—19 tok/s3.43 s
cc-glm-5.1Takes text, returns text.200K131K$0.06$0.22/M———
coding-glm-4.6Takes text, returns text.200K131K$0.06$0.22/M$0.011/M33 tok/s3.77 s
coding-glm-4.7Takes text, returns text.200K131K$0.06$0.22/M$0.011/M38 tok/s1.75 s
coding-glm-5Takes text, returns text.200K131K$0.06$0.22/M—52 tok/s2.49 s
coding-glm-5.1Takes text, returns text.200K131K$0.06$0.22/M—52 tok/s3.41 s
glm-4.7Takes text, returns text.200K131K$0.274$1.0959/M$0.0548/M20 tok/s5.66 s
glm-5v-turboTakes text, vision, video, returns text.200K131K$0.7042$3.0985/M$0.169/M23 tok/s2.07 s
glm-5.1Takes text, returns text.200K131K$0.845$3.38/M$0.1831/M22 tok/s1.26 s
coding-glm-4.5-airTakes text. Output modality not published.131K—$0.014$0.084/M—22 tok/s7.71 s
glm-4.6vTakes text, vision, video, returns text.131K33K$0.137$0.411/M$0.0274/M——
glm-4.5-airTakes text. Output modality not published.131K98K$0.14$0.84/M—91 tok/s1.42 s
glm-4.5Takes text. Output modality not published.131K98K$0.4$1.6/M—107 tok/s0.46 s
glm-4.5vTakes text, vision, video, returns text.66K16K$0.274$0.822/M—81 tok/s8.45 s
glm-ocrTakes vision, returns text.32K—$0.0282$0.0282/M———
embedding-2Takes text. Output modality not published.8K—$0.0686$0.0686/M———
embedding-3Takes text. Output modality not published.8K—$0.0686$0.0686/M———
glm-imageTakes text, returns vision.——FreeFree/M———
Pro/THUDM/GLM-4.1V-9B-Thinking——$0.04$0.16/M———
THUDM/GLM-4-9B-0414——$0.05$0.05/M———
THUDM/GLM-Z1-9B-0414——$0.05$0.05/M———
cc-glm-4.6——$0.06$0.22/M———
cc-glm-4.7——$0.06$0.22/M———
THUDM/GLM-4-32B-0414——$0.08$0.08/M———
THUDM/GLM-Z1-32B-0414——$0.08$0.08/M———
glm-4-flash——$0.1$0.1/M———
THUDM/GLM-4.1V-9B-Thinking——$0.1$0.1/M———
doubao-1-5-pro-32k-250115——$0.108$0.27/M———
chatglm_lite——$0.2858$0.2858/M———
alicloud-glm-4.7——$0.411$1.9178/M$0.411/M44 tok/s1.01 s
alicloud-glm-5——$0.5634$2.5353/M$0.1127/M49 tok/s1.78 s
doubao-1-5-pro-256k-250115——$0.684$1.2312/M———
glm-3-turbo——$0.71$0.71/M———
chatglm_std——$0.7144$0.7144/M———
chatglm_turbo——$0.7144$0.7144/M———
glm-4.5-airxTakes text. Output modality not published.——$1.1$4.51/M$0.22/M——
chatglm_pro——$1.4286$1.4286/M———
glm-4v-plus——$2$2/M———
glm-zero-preview——$2$2/M———
glm-4.5-xTakes text. Output modality not published.——$2.2$8.91/M$0.44/M1 tok/s0.59 s
cbs-glm-4.7——$2.25$2.75/M———
glm-4-plus——$8$8/M———
cogview-3-plus——$10$10/M———
glm-4——$14.2$14.2/M———
glm-4v——$14.2$14.2/M———
code-davinci-edit-001——$20$20/M———
cogview-3——$35.5$35.5/M———

Prices are USD per million tokens; cache read and cache write are the rates for prompt-cache hits and for writing a prompt into the cache. Throughput and latency are measured on AIHubMix — the same figures the model detail page shows — not vendor claims. A dash means the catalog does not publish that field for that model, which is not the same as the model not supporting it.

Z.AI on AIHubMix

Which Z.AI model should I start with?

coding-glm-4.6-free is free on input — the cheapest entry here that declares tool calling, and it carries a 200K context. Move up to cogview-3 when answer quality matters more than cost, or to coding-glm-5.3-free for long-form reasoning.

Which of these models reason before answering?

29 of the 71 models here declare a reasoning phase — they work through the problem before producing an answer, which helps on multi-step problems at the cost of extra output tokens. Use the Reasoning filter above the table to see them. The catalog does not record anything further about how they differ, so this page does not sort them into families.

Why are there several entries for the same model?

Because each row is a route you can call, not a model release. Some IDs name an upstream (azure-, alicloud-, cc-), some are the open-weight repository form (THUDM/…), and some differ only in capitalisation, kept so older integrations keep working.

The catalog does not carry a field saying which of those a given row is, so this page does not sort them into buckets it would have to invent. Every row shows that route’s own price, context and speed — compare those directly, and open a model to see the upstreams that serve it.

How is cached input billed?

The Cache read column is the rate for input tokens served from the prompt cache — for example coding-glm-4.6 bills cache hits at 18.33% of the input rate and coding-glm-4.7 bills cache hits at 18.33% of the input rate. Cache write is the surcharge for putting a prompt into the cache in the first place, and only a few upstreams bill it separately. A dash in either column means the catalog carries no cache rate for that model, so plan on paying the full input rate.

Do I need a separate Z.AI account?

No. One AIHubMix key covers every model on this page, and switching between them is a change to the model string — billing, rate limits, and logs stay in one place.

Start calling Z.AI in one line

One key, one endpoint, 911 models across 42 model authors.