Qwen2-VL 7B Instruct

Permissive

Alibaba · Qwen · Released 2024-08

7BMultimodal

A vision-language model that handles arbitrary image resolutions and understands video, strong on document/OCR-style tasks for its size.

Strengths

  • +Apache 2.0 license
  • +Native dynamic-resolution image understanding
  • +Video understanding in addition to images

Limitations

  • -Higher-resolution inputs increase inference cost/latency
  • -Smaller variant trails Qwen2-VL-72B on fine-grained visual reasoning

License

Apache 2.0

use commercially with attribution niceties

Hardware

runs on a good consumer GPU (or Apple Silicon) with quantization

Links

Stats

Qwen/Qwen2-VL-7B-Instruct— downloads·— likes

via Hugging Face

Related models