GLM-4.6: hardware requirements
GLM-4.6 needs about 220 GB of memory at Q4 quantization — a 215 GB download plus room for context and overhead. That is beyond consumer hardware; use it through an API instead.
Parameters
357B
Tier
Server only
Licence
See model card
Released
2025-09-30
Will it run on your machine?
8GB RAMNot enough memory16GB RAMNot enough memory24GB RAMNot enough memory32GB RAMNot enough memory64GB RAMNot enough memory128GB RAMNot enough memory8GB VRAMNot enough memory12GB VRAMNot enough memory16GB VRAMNot enough memory24GB VRAMNot enough memory32GB VRAMNot enough memory
Download sizes by quantization
Lower quantization means a smaller file and less memory, at some cost to quality. Q4_K_M is the usual starting point. What is this?
| Quant | Download | RAM @ 4K | RAM @ 32K | Source |
|---|---|---|---|---|
| Q4_K_Mrecommended | 215 GB | 220 GB | 231 GB | estimated |
| Q5_K_M | 253 GB | 258 GB | 268 GB | estimated |
| Q6_K | 293 GB | 297 GB | 308 GB | estimated |
| Q8_0 | 379 GB | 384 GB | 395 GB | estimated |
| BF16 | 714 GB | 718 GB | 729 GB | estimated |
“Measured” sizes are read from published GGUF files. “Estimated” sizes are derived from the parameter count and are typically within a few percent. The RAM columns add the context cache and runtime overhead to the download size.
Other sizes of GLM-4.6
Model card on HuggingFace →API specs & pricing →Verified 2026-09-13Tracked for 12 months after release