Llama 2 13B Chat AWQ is an open-source language model by TheBloke. Features: 13b LLM, VRAM: 7.2GB, Context: 4K, License: llama2, Quantized, LLM Explorer Score: 0.09.
Llama 2 13B Chat AWQ Parameters and Internals
Model Type
Use Cases
Areas:
Applications: dialogue, natural language generation
Primary Use Cases: assistant-like chat, natural language generation tasks
Limitations: use only in English, avoid inappropriate legal applications
Considerations: Use proper formatting for optimal feature extraction and performance.
Additional Notes AWQ models support efficient inference and faster execution. Compatible with vLLM and AutoAWQ.
Training Details
Data Sources: publicly available online data
Data Volume:
Methodology: auto-regressive language model with transformer architecture, supervised fine-tuning (SFT), and reinforcement learning with human feedback (RLHF)
Context Length:
Hardware Used: A100-80GB GPUs (TDP of 350-400W)
Model Architecture:
Safety Evaluation
Methodologies: truthfulQA, Toxigen benchmarks
Findings: Llama-2-Chat models demonstrate high-performance in safety tests
Risk Categories:
Ethical Considerations: Testing conducted to date has been in English and may not cover all scenarios. Potential for inaccurate or objectionable responses exists.
Responsible Ai Considerations
Fairness: Testing conducted to date has been in English. Model may exhibit bias.
Transparency: Limited model transparency as it is fine-tuned and pretrained with human feedback.
Accountability: Meta is accountable for the model's performance and outputs.
Mitigation Strategies: Developers advised to perform safety testing and tuning tailored to applications.
Input Output
Input Format: Structured text with prompt template
Accepted Modalities:
Output Format:
Performance Tips: Use 'INST' for better performance in chat tasks.
LLM Name Llama 2 13B Chat AWQ Repository π€ https://huggingface.co/TheBloke/Llama-2-13B-chat-AWQ Model Name Llama 2 13B Chat Model Creator Meta Llama 2 Base Model(s) Llama 2 13B Chat Hf meta-llama/Llama-2-13b-chat-hf Model Size 13b Required VRAM 7.2 GB Updated 2026-07-15 Maintainer TheBloke Model Type llama Model Files 7.2 GB Supported Languages en AWQ Quantization Yes Quantization Type awq Model Architecture LlamaForCausalLM License llama2 Context Length 4096 Model Max Length 4096 Transformers Version 4.31.0.dev0 Tokenizer Class LlamaTokenizer Beginning of Sentence Token <s> End of Sentence Token </s> Unk Token <unk> Vocabulary Size 32000 Torch Data Type float16
Rank the Llama 2 13B Chat AWQ capabilities
Have you tried this model? Rate its performance β this feedback helps the ML community find the right model for their needs.
Instruction Following and Task Automation
Factuality and Completeness of Knowledge
Censorship and Alignment
Data Analysis and Insight Generation
Text Generation
Text Summarization and Feature Extraction
Code Generation
Multi-Language Support and Translation
Best Alternatives to Llama 2 13B Chat AWQ
Note: green Score (e.g. "73.2 ") means that the model is better than TheBloke/Llama-2-13B-chat-AWQ .
Expand