GLM-4.7-Flash
zai-org · moe · 31.2B parameters · 202,752 context
Parameters
31.2B
Context Window
198K tokens
Architecture
MoE
Best GPU
H200 SXM
Intelligence Brief
GLM-4.7-Flash is a 31.2B parameter Mixture-of-Experts (64 experts, 4 active) model from zai-org, featuring Multi-Head Attention (MHA) with 47 layers and 2,048 hidden dimensions. With a 202,752 token context window, it supports tools, structured output, code, math, multilingual, reasoning. For self-hosted inference, H200 SXM delivers optimal throughput at $2553/month.
Recent changes
Loading…
Related models
4 suggestions
GLM-4 9BGLM-4 · 9.4B$0.150/M out
MiniMax-M2.1MiniMax-M2.1 · 7B—
MiniMax-M2.5MiniMax-M2 · 7B$1.20/M out
MiniMax-M2MiniMax-M2 · 7B$1.10/M out
Picks: same family first, then same vendor within ±2× params, then top tag-overlap matches. Price shown is the cheapest Output $/M across providers — the row's page shows the canonical anchor.
Architecture Details
Memory Requirements
BF16 Weights
62.4 GB
FP8 Weights
31.2 GB
INT4 Weights
15.6 GB
GPU Compatibility Matrix
GLM-4.7-Flash is compatible with 62% of GPU configurations across 41 GPUs at 3 precision levels.
GPU Recommendations
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
1.1K tok/s
Latency (ITL)
1.0ms
Est. TTFT
0ms
Cost/Month
$2553
Cost/M Tokens
$0.93
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
1.1K tok/s
Latency (ITL)
1.0ms
Est. TTFT
0ms
Cost/Month
$1794
Cost/M Tokens
$0.65
FP8 · 1 GPU · tensorrt-llm
100/100
score
Throughput
1.1K tok/s
Latency (ITL)
1.0ms
Est. TTFT
0ms
Cost/Month
$1794
Cost/M Tokens
$0.65
Deployment Options
API Deployment
No API pricing available
Single GPU
H200 SXM
$2553/mo
Min VRAM: 31 GB
Multi-GPU
RTX A6000 x2
857.4 tok/s
TP· $930/mo
API Pricing Comparison
No API pricing data available for this model.
Performance Estimates
Throughput by GPU
VRAM Breakdown (H200 SXM, FP8)
Precision Impact
bf16
62.4 GB
weights/GPU
fp8
31.2 GB
weights/GPU
~1.1K tok/s
Capabilities
Features
Supported Frameworks
Supported Precisions
Where to Deploy GLM-4.7-Flash
Self-Hosted Infrastructure
Similar Models
GLM-4 9B
9.4B params · dense
Quality: 50
from $0.15/M
Gemma 4 31B-IT
31B params · dense
Quality: 77
from $0.00/M
Qwen 3 30B-A3B
30.5B params · moe
Quality: 70
from $0.45/M
Claude Haiku 4.5
30B params · moe
Quality: 50
from $5.00/M
JAIS 30B
30B params · dense
Quality: 50
Frequently Asked Questions
How much VRAM does GLM-4.7-Flash need for inference?
GLM-4.7-Flash requires approximately 62.4 GB of VRAM at BF16 precision, 31.2 GB at FP8, or 15.6 GB at INT4 quantization. Additional VRAM is needed for KV-cache (721920 bytes per token) and activations (~0.00 GB).
What is the best GPU for GLM-4.7-Flash?
The top recommended GPU for GLM-4.7-Flash is the H200 SXM using FP8 precision. It achieves approximately 1.1K tokens/sec at an estimated cost of $2553/month ($0.93/M tokens). Score: 100/100.
How much does GLM-4.7-Flash inference cost?
GLM-4.7-Flash inference costs vary by provider and GPU setup. Use our calculator for detailed cost estimates across all providers.