Qwen1.5 72B Chat GPTQ is an open-source language model by LoneStriker. Features: 72b LLM, VRAM: 45.4GB, Context: 32K, License: other, Quantized, LLM Explorer Score: 0.24, ELO: 1233, Arc: 68.5, HellaSwag: 86.4, MMLU: 77.4, GSM8K: 20.4.
Qwen1.5 72B Chat GPTQ Benchmarks
nn.n% — How the model compares to the reference models: Anthropic Sonnet 3.5 ("so35"), GPT-4o ("gpt4o") or GPT-4 ("gpt4").
Qwen1.5 72B Chat GPTQ Parameters and Internals
| Model Type | | text generation, chat model |
|
| Additional Notes | | The beta version does not include GQA and the mixture of SWA and full attention. DPO improves human preference but lowers benchmark evaluation. |
|
| Supported Languages | | en (multilingual capabilities include English) |
|
| Training Details |
| Methodology: | | Supervised finetuning and direct preference optimization (DPO) |
|
| Context Length: | |
| Model Architecture: | | Transformer architecture with SwiGLU activation, attention QKV bias, group query attention, mixture of sliding window attention and full attention |
|
|
| Input Output |
| Accepted Modalities: | |
| Performance Tips: | | Use provided hyper-parameters in 'generation_config.json' for optimal performance. |
|
|
| LLM Name | Qwen1.5 72B Chat GPTQ |
| Repository π€ | https://huggingface.co/LoneStriker/Qwen1.5-72B-Chat-GPTQ |
| Base Model(s) | Qwen/Qwen1.5-72B-Chat Qwen/Qwen1.5-72B-Chat |
| Model Size | 72b |
| Required VRAM | 45.4 GB |
| Updated | 2026-07-28 |
| Maintainer | LoneStriker |
| Model Type | qwen2 |
| Model Files | 45.4 GB |
| Supported Languages | en |
| GPTQ Quantization | Yes |
| Quantization Type | gptq|4bit |
| Model Architecture | Qwen2ForCausalLM |
| License | other |
| Context Length | 32768 |
| Model Max Length | 32768 |
| Transformers Version | 4.37.1 |
| Tokenizer Class | Qwen2Tokenizer |
| Padding Token | <|endoftext|> |
| Vocabulary Size | 152064 |
| Torch Data Type | float16 |
| Errors | replace |
Best Alternatives to Qwen1.5 72B Chat GPTQ