Qwen2 VL 2B Instruct GPTQ Int4 is an open-source language model by Qwen. Features: 2b LLM, VRAM: 2.5GB, Context: 32K, License: apache-2.0, Quantized, Instruction-Based, LLM Explorer Score: 0.15.
visual understanding, video-based QA, mobile and robotic integrations
Primary Use Cases:
visual question answering, dialog, content creation, multilingual support
Limitations:
Lack of audio support, Updates up to June 2023, Limited individual/IP recognition, Limited complex instruction handling, Low counting accuracy, Weak spatial reasoning
Additional Notes
Reports quantized model performance across various tasks, highlighting strengths in multimodal integration and weaknesses, such as limitations in audio and complex reasoning.
Supported Languages
English (native), Chinese (native), European languages (high), Japanese (high), Korean (high), Arabic (high), Vietnamese (high)
Training Details
Data Volume:
up to June 2023
Model Architecture:
Naive Dynamic Resolution, Multimodal Rotary Position Embedding (M-ROPE)
Input Output
Input Format:
Images, Videos (local files, base64, URLs)
Accepted Modalities:
text, image
Output Format:
Textual descriptions
Performance Tips:
Enabling flash_attention_2 recommended for better acceleration and memory saving.
Release Notes
Version:
Qwen2-VL-2B-Instruct-GPTQ-Int4
Date:
2023
Notes:
Quantized model version with multi-language support and enhanced image and video processing capabilities.