Skip to content
Tier A — Frontier
Runs in:CNMade in:China
Z.ai (GLM / Zhipu)

GLM-4.6V (vision)

Tier A — Frontier · 205K tokens

Tokonomix Editorial Team·Reviewed by Mes Kalkan··

GLM-4.6V is the vision-capable model of the GLM-4.6 line: it accepts images alongside text and reasons over both. It pairs multimodal understanding with the GLM-4.6 line’s large context window, at a notably low price.

Section 01

Speed analysis

Latency measured across all benchmark runs. P50 (median) and P95 (95th percentile) give a realistic picture of response speed under normal and peak load.

P50 latency (median)P95 latency56 runs
821849116161238303150007-0807-22ms
Section 02

Quality scores

Evaluation results from judge-model scoring across diverse task categories. Scores reflect coherence, accuracy and instruction-following.

75
Coding
45
Creative
Section 03

Pricing history

Direct provider rates per million tokens, plus a typical-conversation cost estimate.

💰
API rates — GLM-4.6V (vision)
$0.3000 per 1M input tokens
$0.9000 per 1M output tokens
≈ $0.0004 per typical conversation (800 tokens)
Input vs output price (per 1M tokens)
per 1M input tokens$0.3000
per 1M output tokens$0.9000

Pricing over time

Input & output per 1M tokens · step-line = price changes

$0.3000

input / 1M

— stable

$0.9000

output / 1M

— stable

2026-07-122026-07-192026-07-19
Input
Output
Price change
⟳ synced weekly
Section 04

Tokens per second

Throughput in tokens per second, derived from measured P50 latency. Higher is better; fluctuations track provider-side load.

Throughput (tokens / s)96 / avg 140
2415

Estimated from P50 latency × 200 output tokens — the absolute number depends on this assumption; the trend is what matters.

Section 05

Capabilities

jsonnotes: GLM emits a non-standard reasoning_content field beside content; read content for the answer.toolsvision
Section 06

Availability

Availability

How often this model answers when we call it — measured across real API requests and live tests over the last 30 days. This is separate from quality: these numbers only tell you whether the model responds, not how good the answer is.

Last 7 days

Last 30 days

100.0%

n=1

Median response time

3,997ms

n=1

Based on 175 measurements over the last 30 days.

Technical details

Only live API calls and live-test requests count — internal probes and benchmark runs are excluded.

Calls with a custom API key (BYOK) are excluded: those failures are key-specific, not a sign of model downtime.

Failed calls are NOT included in quality scores — quality is measured on successful responses only. Availability and quality are independent signals.

Median response time (p50) across successful calls with a recorded duration. Outliers (very slow or very fast calls) pull the median less than the average.

Total calls (30d)

1

OK responses (30d)

1

Total calls (7d)

0

OK responses (7d)

0

Section 07

Tokonomix benchmark verdicts

⚖️
Endorsed by 1 judge
Independent LLM judges evaluated this model on our weekly intelligence tests
claude-sonnet-4-561/100 · 6 runs
3 correct0 partial3 wrong50% accuracy
2026-07-19

GLM-4.6V debuts with vision capabilities and multimodal support

GLM-4.6V enters the benchmark landscape as a new multimodal model from Zhipu AI, bringing vision capabilities alongside text processing. The model demonstrates solid performance across vision tasks, achieving 49.0% on MMMU and 71.6% on MathVista, placing it in the competitive middle tier for multimodal understanding. On text-only benchmarks, it scores 56.3% on MMLU and 49.8% on GPQA Diamond, showing respectable but not leading-edge performance. The model supports both JSON output and tool use, making it suitable for structured applications and agentic workflows. Notably, it handles Chinese language tasks well at 71.8% on MMMLU, reflecting its origins from a Chinese research team. The addition of vision capabilities expands the GLM series into multimodal territory, though performance metrics suggest it targets practical applications rather than state-of-the-art benchmarks. Users seeking a capable multimodal model with good Chinese language support and standard API features will find GLM-4.6V a workable option, though those requiring top-tier performance on complex reasoning or vision tasks may need to consider alternatives.

Quality

Latency p50

Test runs

0

Vision capabilities added JSON and tool support Strong Chinese language performance
Section 08

Full model profile

GLM-4.6V: long-context vision from Zhipu

GLM-4.6V is the vision-capable model of the GLM-4.6 line: it accepts images alongside text and reasons over both. It pairs multimodal understanding with the GLM-4.6 line’s large context window, at a notably low price.

z.ai publishes GLM-4.6V at $0.30 per 1M input tokens and $0.90 per 1M output tokens — an aggressively low price for a vision-capable model.

It advertises a large ~200K-token context window, so images can be combined with substantial text in one call.

Architecture & training signals

GLM-4.6V is the vision variant of Zhipu AI’s GLM-4.6. It extends the text model with image input, producing text output that reasons over the supplied images. It exposes tool-calling and JSON over an OpenAI-compatible endpoint. Its output modality is text (it describes/reasons about images, it does not generate them). Like the rest of the GLM line, it returns a non-standard reasoning_content field alongside content in its OpenAI-compatible responses; integrations should read content for the final answer and treat reasoning_content as an optional trace.

Where it shines

  • Image understanding — describing, comparing, extracting information from pictures and screenshots.
  • Combining images with long text context in a single call.
  • Very low price for a multimodal model.

Where it falls short

  • It reads images, it does not generate them — for image generation use z.ai’s GLM-Image / CogView models.
  • No Tokonomix benchmark data yet; verify vision quality on your own images.
  • Non-EU hosting.

Real-world use cases

  • Document, chart and screenshot understanding.
  • Visual QA and multimodal analysis at scale.
  • A cheap vision proposer in a multimodal consensus panel.

Tokonomix benchmark snapshot

GLM-4.6V is newly registered on Tokonomix and not yet activated, so we have not run it through our weekly intelligence test or speed benchmark. There are no Tokonomix scores to report yet — and we will not invent any.

When it goes live, it enters the same weekly harness as every other model: identical prompts, an independent cross-family judge, and reproducible latency and cost measurements. Until then, treat the pricing and capability notes on this page as the vendor-published starting point, not as measured Tokonomix results.

EU privacy & data residency

GLM-4.6V is built by Zhipu AI (z.ai), a China-headquartered lab, and is served from non-EU infrastructure. This is important to state plainly: routing a prompt to this model is not an EU-data-residency or GDPR-sovereign choice, and Tokonomix will never tag it as one.

If your use case requires data to stay within the EU, pick a model whose provider is EU-hosted (for example our OVH or Azure-EU routes) rather than a GLM model. Tokonomix keeps z.ai out of every EU-only / sovereign routing set by design. Use GLM where its capability or price is the priority and cross-border processing is acceptable for that workload.

Verdict & alternatives

GLM-4.6V is a strong-value vision model: multimodal understanding plus long context at a low price. For a free vision option, GLM-4.6V Flash; for image generation (not understanding), the GLM-Image / CogView models.

Last automated test
Jul 22, 2026 · 02:04 UTC · Speed benchmark
P50 latency
2085 ms
P95 latency
2196 ms
Errors
0 / 6 runs
Last reviewed by Tokonomix Team·July 8, 2026