GLM 4.5 Vision
Z.AI logo

GLM 4.5 Vision

glm-4.5vllms.txt
Z.AI
GLM-4.5V is a vision-language foundational model designed for multimodal agent applications. Based on a mixture-of-experts (MoE) architecture, it has 106 billion parameters and 12 billion active parameters. It delivers outstanding performance in video understanding, image question answering, OCR, and document parsing, and achieves significant improvements in front-end web encoding, basic reasoning, and spatial reasoning.

Pricing

TierPricingCache Read
Input<=32K
$0.274$1.096
-
32K<Input<=64K
$0.548$1.644
-

Input Modalities

  • Text
  • Vision
  • Video

Output Modalities

  • Text

Providers

Z.AI zai-glm-4.5v
Pricing$0.274$0.822
Pricing$0.548$1.644
Context64K
Max output16K
Latency8.4S
Throughput81.2TPS
Uptime
100.00% uptime 3 days ago
100.00% uptime 2 days ago
0.00% uptime yesterday

Performance for glm-4.5v

Uptime is the percentage of requests that succeeded over the past 72 hours. AIHubMix continuously monitors every provider and automatically retries with the next-best provider when one returns an error or responds too slowly; Latency is total round-trip time (lower is better); Throughput is how fast the model writes (tokens per second, higher is better).

Uptime
Loading...
Latency
Loading...
Throughput
Loading...

Try this model

Python
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["AIHUBMIX_API_KEY"],
    base_url="https://aihubmix.com/v1",
)

response = client.chat.completions.create(
    model="glm-4.5v",
    messages=[
      {
        "role": "user",
        "content": "Hello, how are you?"
      }
    ],
    max_tokens=1024,
    stream=False,
)

print(response.choices[0].message.content)

Frequently asked questions

What is GLM 4.5 Vision?

GLM-4.5V is a vision-language foundational model designed for multimodal agent applications. Based on a mixture-of-experts (MoE) architecture, it has 106 billion parameters and 12 billion active parameters. It delivers outstanding performance in video understanding, image question answering, OCR, and document parsing, and achieves significant improvements in front-end web encoding, basic reasoning, and spatial reasoning.