Qwen2-VL 7B Instruct
PermissiveAlibaba · Qwen · Released 2024-08
7BMultimodal
A vision-language model that handles arbitrary image resolutions and understands video, strong on document/OCR-style tasks for its size.
Strengths
- +Apache 2.0 license
- +Native dynamic-resolution image understanding
- +Video understanding in addition to images
Limitations
- -Higher-resolution inputs increase inference cost/latency
- -Smaller variant trails Qwen2-VL-72B on fine-grained visual reasoning
License
Apache 2.0
use commercially with attribution niceties
Hardware
runs on a good consumer GPU (or Apple Silicon) with quantization
Links
Stats
Qwen/Qwen2-VL-7B-Instruct— downloads·— likes
via Hugging Face