LLM EXPLORER 57,252 MODELS INDEXED

GLM 4.7 Flash AWQ 4bit by cyankiwi

By cyankiwi · 489953 downloads

GLM 4.7 Flash AWQ 4bit is an open-source language model by cyankiwi. Features: 32.1b LLM, VRAM: 20.3GB, Context: 198K, License: mit, MoE, Quantized, LLM Explorer Score: 0.47, ELO: 1368.

  Arxiv:2508.06471   4bit   Awq Base model:quantized:zai-org/g... Base model:zai-org/glm-4.7-fla...   Compressed-tensors   Conversational   Deploy:azure   En   Endpoints compatible   Glm4 moe lite   Moe   Quantized   Region:us   Safetensors   Sharded   Tensorflow   Zh

GLM 4.7 Flash AWQ 4bit Benchmarks

nn.n% — How the model compares to the reference models: Anthropic Sonnet 3.5 ("so35"), GPT-4o ("gpt4o") or GPT-4 ("gpt4").

GLM 4.7 Flash AWQ 4bit Parameters and Internals

LLM NameGLM 4.7 Flash AWQ 4bit
Repository πŸ€—https://huggingface.co/cyankiwi/GLM-4.7-Flash-AWQ-4bit 
Base Model(s)  zai-org/GLM-4.7-Flash   zai-org/GLM-4.7-Flash
Model Size32.1b
Required VRAM20.3 GB
Updated2026-08-09
Maintainercyankiwi
Model Typeglm4_moe_lite
Model Files  5.4 GB: 1-of-4   5.4 GB: 2-of-4   5.4 GB: 3-of-4   4.1 GB: 4-of-4
Supported Languagesen zh
AWQ QuantizationYes
Quantization Typeawq|4bit
Model ArchitectureGlm4MoeLiteForCausalLM
Licensemit
Context Length202752
Model Max Length202752
Transformers Version5.3.0.dev0
Tokenizer ClassTokenizersBackend
Padding Token<|endoftext|>
Vocabulary Size154880