LLaVA-1.6 Mistral 7B
PermissiveLLaVA community · LLaVA · Released 2024-01
7BMultimodal
One of the earliest and still widely-used open vision-language models, combining a CLIP vision encoder with a Mistral 7B backbone.
Strengths
- +Apache 2.0 (via Mistral 7B base) license
- +Large ecosystem of tooling and derivative fine-tunes
- +Runs on a single consumer GPU
Limitations
- -Fixed, lower input image resolution than newer VLMs like Qwen2-VL
- -OCR and fine-grained document understanding trail 2024+ models
License
Apache 2.0
use commercially with attribution niceties
Hardware
runs on a good consumer GPU (or Apple Silicon) with quantization
Links
Stats
llava-hf/llava-v1.6-mistral-7b-hf— downloads·— likes
via Hugging Face