LLM EXPLORER 60,702 MODELS INDEXED

LLaDA2.0 Flash 4bit by mlx-community

By mlx-community · 18 downloads

LLaDA2.0 Flash 4bit is an open-source language model by mlx-community. Features: 102.9b LLM, VRAM: 57.8GB, Context: 32K, License: apache-2.0, Quantized, LLM Explorer Score: 0.21.

  4-bit   4bit Base model:inclusionai/llada2.... Base model:quantized:inclusion...   Conversational   Custom code   Diffusion   Dllm   Llada2 moe   Mlx   Quantized   Region:us   Safetensors   Sharded   Tensorflow   Text generation

LLaDA2.0 Flash 4bit Parameters and Internals

LLM NameLLaDA2.0 Flash 4bit
Repository πŸ€—https://huggingface.co/mlx-community/LLaDA2.0-flash-4bit 
Base Model(s)  inclusionAI/LLaDA2.0-flash   inclusionAI/LLaDA2.0-flash
Model Size102.9b
Required VRAM57.8 GB
Updated2026-07-18
Maintainermlx-community
Model Typellada2_moe
Model Files  5.4 GB: 1-of-12   4.9 GB: 2-of-12   4.9 GB: 3-of-12   4.9 GB: 4-of-12   4.9 GB: 5-of-12   4.9 GB: 6-of-12   4.9 GB: 7-of-12   4.9 GB: 8-of-12   4.9 GB: 9-of-12   4.9 GB: 10-of-12   4.9 GB: 11-of-12   3.4 GB: 12-of-12
Quantization Type4bit
Model ArchitectureLLaDA2MoeModelLM
Licenseapache-2.0
Context Length32768
Model Max Length32768
Transformers Version4.51.0
Tokenizer ClassPreTrainedTokenizerFast
Padding Token<|endoftext|>
Vocabulary Size157184
Torch Data Typebfloat16

Best Alternatives to LLaDA2.0 Flash 4bit

Best Alternatives
Context / RAM
Downloads
Likes
LLaDA2.2 Flash OptiQ 2bit128K / 38.9 GB613
LLaDA2.0 Flash 8bit32K / 109.3 GB111
LLaDA2.0 Flash Preview 4bit16K / 57.8 GB123
LLaDA2.2 Flash128K / 204.5 GB32857
LLaDA2.1 Flash32K / 204.5 GB2982093
LLaDA2.0 Flash32K / 206 GB51869
LLaDA2.0 Flash CAP32K / 206 GB659
LLaDA2.0 Flash Preview16K / 205.9 GB2467
Note: green Score (e.g. "73.2") means that the model is better than mlx-community/LLaDA2.0-flash-4bit.