BLIP-3 / xGen-MM
Salesforce's multimodal model for enterprise vision tasks
Verdict
Strong at visual grounding, image captioning, and structured extraction. Good for enterprise workflows requiring custom vision-language pipelines. Apache 2.0 licensed.
Other Multimodal & Vision
- LLaVA-NeXTStable
Leading open-source vision-language model with strong reasoning
- Pixtral LargeStable
Mistral's 124B vision-language model with 128K context
- InternVL3Stable
Open-source VLM rivalling closed frontier models on vision benchmarks
- Qwen-VL-MaxStable
Alibaba's flagship vision-language model with video understanding
- Florence-2Production
Microsoft's unified vision foundation model for multiple tasks
- MolmoExperimental
Allen AI's fully open VLM with pointing and grounding

