LLM EXPLORER 59,358 MODELS INDEXED

Llama 2 13B Chat Hf 4bit G64 HQQ by mobiuslabsgmbh

By mobiuslabsgmbh · 4 downloads

Llama 2 13B Chat Hf 4bit G64 HQQ is an open-source language model by mobiuslabsgmbh. Features: 13b LLM, VRAM: 7.6GB, Context: 4K, License: llama2, Quantized, LLM Explorer Score: 0.1.

  4bit   Conversational   Llama   Quantized   Region:us

Llama 2 13B Chat Hf 4bit G64 HQQ Parameters and Internals

Model Type 
text generation
Use Cases 
Primary Use Cases:
Chatbot applications
Limitations:
Only supports single GPU runtime., Not compatible with HuggingFace's PEFT.
Additional Notes 
This model is a quantized version of Llama-2-13B-chat, optimized for efficient usage with reduced precision.
Input Output 
Input Format:
PyTorch tensors after tokenization.
Accepted Modalities:
text
Output Format:
Generated text from processed prompt.
LLM NameLlama 2 13B Chat Hf 4bit G64 HQQ
Repository πŸ€—https://huggingface.co/mobiuslabsgmbh/Llama-2-13b-chat-hf-4bit_g64-HQQ 
Model Size13b
Required VRAM7.6 GB
Updated2026-07-07
Maintainermobiuslabsgmbh
Model Typellama
Model Files  7.6 GB
Quantization Type4bit
Model ArchitectureLlamaForCausalLM
Licensellama2
Context Length4096
Model Max Length4096
Transformers Version4.35.2
Tokenizer ClassLlamaTokenizer
Beginning of Sentence Token<s>
End of Sentence Token</s>
Unk Token<unk>
Vocabulary Size32000
Torch Data Typefloat16

Best Alternatives to Llama 2 13B Chat Hf 4bit G64 HQQ

Best Alternatives
Context / RAM
Downloads
Likes
Llama13b 32K Illumeet Finetune32K / 26 GB50
...Maid V3 13B 32K 8.0bpw H8 EXL232K / 13.2 GB81
...Maid V3 13B 32K 6.0bpw H6 EXL232K / 10 GB51
WhiteRabbitNeo 13B V116K / 26 GB3146453
CodeLlama 13B Python Fp1616K / 26 GB12125
CodeLlama 13B Fp1616K / 26 GB967
CodeLlama 13B Instruct Fp1616K / 26 GB13728
Codellama 13B Bnb 4bit16K / 7.2 GB1295
...Llama 13B Instruct Hf 4bit MLX16K / 7.8 GB1133
WhiteRabbitNeo 13B V1 4bit Mlx16K / 7.8 GB1472
Note: green Score (e.g. "73.2") means that the model is better than mobiuslabsgmbh/Llama-2-13b-chat-hf-4bit_g64-HQQ.