LLM EXPLORER 60,702 MODELS INDEXED

DeepSeek V4 Flash W4A16 FP8 by canada-quant

By canada-quant · 2621 downloads

DeepSeek V4 Flash W4A16 FP8 is an open-source language model by canada-quant. Features: 44.1b LLM, VRAM: 152.5GB, Context: 1024K, License: mit, LLM Explorer Score: 0.28.

Base model:deepseek-ai/deepsee... Base model:quantized:deepseek-...   Compressed-tensors   Deepseek   Deepseek v4   En   Fp8   Gptq   Mixture-of-experts   Moe   Region:us   Safetensors   Sharded   Tensorflow   Vllm   W4a16   Zh

DeepSeek V4 Flash W4A16 FP8 Parameters and Internals

LLM NameDeepSeek V4 Flash W4A16 FP8
Repository 🤗https://huggingface.co/canada-quant/DeepSeek-V4-Flash-W4A16-FP8 
Base Model(s)  DeepSeek V4 Flash   deepseek-ai/DeepSeek-V4-Flash
Model Size44.1b
Required VRAM152.5 GB
Updated2026-08-04
Maintainercanada-quant
Model Typedeepseek_v4
Model Files  50.0 GB: 1-of-4   50.0 GB: 2-of-4   50.0 GB: 3-of-4   2.5 GB: 4-of-4
Supported Languagesen zh
Model ArchitectureDeepseekV4ForCausalLM
Licensemit
Context Length1048576
Model Max Length1048576
Transformers Version5.8.0.dev0
Tokenizer ClassTokenizersBackend
Padding Token<|end▁of▁sentence|>
Vocabulary Size129280
Torch Data Typebfloat16

Best Alternatives to DeepSeek V4 Flash W4A16 FP8

Best Alternatives
Context / RAM
Downloads
Likes
DeepSeek V4 Flash W4A16 FP81024K / 152.5 GB55897
Note: green Score (e.g. "73.2") means that the model is better than canada-quant/DeepSeek-V4-Flash-W4A16-FP8.