LLM EXPLORER 63,636 MODELS INDEXED

GLM 4.7 Flash 0.82B MTP by inference-optimization

By inference-optimization · 10 downloads

GLM 4.7 Flash 0.82B MTP is an open-source language model by inference-optimization. Features: 821.3m LLM, VRAM: 1.6GB, Context: 198K, License: mit, MoE, LLM Explorer Score: 0.27.

  Conversational   Endpoints compatible   Glm4 moe lite   Llm-compressor   Moe   Mtp   Pr-3225   Region:us   Safetensors   Sharded   Tensorflow   Test-fixture   Tiny-model

GLM 4.7 Flash 0.82B MTP Parameters and Internals

LLM NameGLM 4.7 Flash 0.82B MTP
Repository πŸ€—https://huggingface.co/inference-optimization/GLM-4.7-Flash-0.82B-MTP 
Model Size821.3m
Required VRAM1.6 GB
Updated2026-10-10
Maintainerinference-optimization
Model Typeglm4_moe_lite
Model Files  0.5 GB: 1-of-4   0.5 GB: 2-of-4   0.5 GB: 3-of-4   0.1 GB: 4-of-4   0.0 GB
Model ArchitectureGlm4MoeLiteForCausalLM
Licensemit
Context Length202752
Model Max Length202752
Transformers Version5.17.0
Tokenizer ClassTokenizersBackend
Padding Token<|endoftext|>
Vocabulary Size154880