LLaVA-1.6 Mistral 7B

Permissive

LLaVA community · LLaVA · Released 2024-01

7BMultimodal

One of the earliest and still widely-used open vision-language models, combining a CLIP vision encoder with a Mistral 7B backbone.

Strengths

  • +Apache 2.0 (via Mistral 7B base) license
  • +Large ecosystem of tooling and derivative fine-tunes
  • +Runs on a single consumer GPU

Limitations

  • -Fixed, lower input image resolution than newer VLMs like Qwen2-VL
  • -OCR and fine-grained document understanding trail 2024+ models

License

Apache 2.0

use commercially with attribution niceties

Hardware

runs on a good consumer GPU (or Apple Silicon) with quantization

Links

Stats

llava-hf/llava-v1.6-mistral-7b-hf— downloads·— likes

via Hugging Face

Related models