LLM EXPLORER 60,384 MODELS INDEXED

ELYZA Japanese Llama 2 13B Fast Instruct 4bit Quantized by SushiTokyo

By SushiTokyo · 6 downloads

ELYZA Japanese Llama 2 13B Fast Instruct 4bit Quantized is an open-source language model by SushiTokyo. Features: 13b LLM, VRAM: 7.8GB, Context: 4K, Quantized, Instruction-Based, LLM Explorer Score: 0.12.

  4-bit   4bit   Endpoints compatible   Gptq   Instruct   Llama   Quantized   Region:us   Safetensors   Sharded   Tensorflow

ELYZA Japanese Llama 2 13B Fast Instruct 4bit Quantized Parameters and Internals

Model Type 
text generation, instruction-following
Additional Notes 
The model is quantized to 4-bit for faster processing. Original model faced difficulty even with RTX4090, but now outputs in about 10 seconds. Accuracy is unmeasured. Refer to quantize.py for quantization code.
LLM NameELYZA Japanese Llama 2 13B Fast Instruct 4bit Quantized
Repository πŸ€—https://huggingface.co/SushiTokyo/ELYZA-japanese-Llama-2-13b-fast-instruct-4bit-quantized 
Model Size13b
Required VRAM7.8 GB
Updated2026-08-06
MaintainerSushiTokyo
Model Typellama
Instruction-BasedYes
Model Files  5.0 GB: 1-of-2   2.8 GB: 2-of-2
Quantization Type4bit
Model ArchitectureLlamaForCausalLM
Context Length4096
Model Max Length4096
Transformers Version4.40.1
Tokenizer ClassLlamaTokenizer
Padding Token</s>
Vocabulary Size44581
Torch Data Typefloat16

Best Alternatives to ELYZA Japanese Llama 2 13B Fast Instruct 4bit Quantized

Best Alternatives
Context / RAM
Downloads
Likes
CodeLlama 13B Instruct Fp1616K / 26 GB13728
...Llama 13B Instruct Hf 4bit MLX16K / 7.8 GB1133
...13B Instruct Nf4 Fp16 Upscaled16K / 26 GB60
Model 007 13b V24K / 26 GB574
...igogne2 Enno 13B Sft Lora 4bit4K / 26 GB7860
Xwin LM 13B V0.2 EXL24K / 5.2 GB123
Mythalion 13B 2.30bpw H4 EXL24K / 4.1 GB43
...lion Kimiko V2.6.05bpw H8 EXL24K / 10.1 GB51
Finance LLM 13B 6.0bpw H6 EXL22K / 10 GB51
Law LLM 13B 4.0bpw H6 EXL22K / 6.8 GB21
Note: green Score (e.g. "73.2") means that the model is better than SushiTokyo/ELYZA-japanese-Llama-2-13b-fast-instruct-4bit-quantized.