LLM EXPLORER 60,828 MODELS INDEXED

Qwen1.5 32B Chat GPTQ Int4 by Qwen

By Qwen · 6616 downloads

Qwen1.5 32B Chat GPTQ Int4 is an open-source language model by Qwen. Features: 32b LLM, VRAM: 19.5GB, Context: 32K, License: other, Quantized, LLM Explorer Score: 0.15.

  Arxiv:2309.16609   4-bit   Chat   Conversational   Deploy:azure   En   Endpoints compatible   Gptq   Quantized   Qwen2   Region:us   Safetensors   Sharded   Tensorflow

Qwen1.5 32B Chat GPTQ Int4 Parameters and Internals

Model Type 
transformer-based, decoder-only, language model
Use Cases 
Areas:
Research, Commercial applications
Applications:
Text generation, Chat models
Primary Use Cases:
Introduction to language models
Limitations:
Code switching
Additional Notes 
Users are warned against deploying this model with vLLM temporarily and are instead advised to use the AWQ model.
Supported Languages 
languages_supported (Multilingual), proficiency_levels (Both base and chat models support multiple languages)
Training Details 
Data Volume:
a large amount of data
Methodology:
pretrained with both supervised finetuning and direct preference optimization
Context Length:
32768
Model Architecture:
Transformer architecture with SwiGLU activation, attention QKV bias, group query attention, mixture of sliding window attention and full attention
Input Output 
Input Format:
Expected structure is a list of message dictionaries containing 'role' and 'content'
Accepted Modalities:
text
Output Format:
Text
Performance Tips:
Use provided hyperparameters in generation_config.json to avoid issues like code switching
LLM NameQwen1.5 32B Chat GPTQ Int4
Repository πŸ€—https://huggingface.co/Qwen/Qwen1.5-32B-Chat-GPTQ-Int4 
Model Size32b
Required VRAM19.5 GB
Updated2026-08-06
MaintainerQwen
Model Typeqwen2
Model Files  4.0 GB: 1-of-5   4.0 GB: 2-of-5   4.0 GB: 3-of-5   4.0 GB: 4-of-5   3.5 GB: 5-of-5
Supported Languagesen
GPTQ QuantizationYes
Quantization Typegptq
Model ArchitectureQwen2ForCausalLM
Licenseother
Context Length32768
Model Max Length32768
Transformers Version4.37.0
Tokenizer ClassQwen2Tokenizer
Padding Token<|endoftext|>
Vocabulary Size152064
Torch Data Typefloat16
Errorsreplace

Best Alternatives to Qwen1.5 32B Chat GPTQ Int4

Best Alternatives
Context / RAM
Downloads
Likes
Qwen2.5 32B Instruct GPTQ Int432K / 19.5 GB47852040
Qwen2.5 32B Instruct GPTQ Int832K / 35.1 GB12158314
...5 Coder 32B Instruct GPTQ Int432K / 19.5 GB2061924
...5 Coder 32B Instruct GPTQ Int832K / 35.1 GB77124
QwQ 32B Preview GPTQ 4bit32K / 16.1 GB3063
Qwen Qwen1.5 32B 4 Bit Gptq32K / 19.2 GB80
Qwen1.5 32B Chat GPTQ Int832K / 34.9 GB101
...Logic Mix 3 32B Mlx OptiQ 5bit128K / 22.3 GB411
...pSeek R1 Distill Qwen 32B 4bit128K / 18.5 GB779551
Hydraulic Deepseek 16bit128K / 65.8 GB520
Note: green Score (e.g. "73.2") means that the model is better than Qwen/Qwen1.5-32B-Chat-GPTQ-Int4.