InternVL3
Open-source VLM rivalling closed frontier models on vision benchmarks
Verdict
One of the strongest open-weight VLMs available, with native multimodal pre-training (rather than bolting vision onto a text-only base) improving reasoning over the 2.5 generation. Strong OCR, chart understanding, and mathematical reasoning from images. Multiple sizes available for different compute budgets.
Other Multimodal & Vision
- LLaVA-NeXTStable
Leading open-source vision-language model with strong reasoning
- Pixtral LargeStable
Mistral's 124B vision-language model with 128K context
- Qwen-VL-MaxStable
Alibaba's flagship vision-language model with video understanding
- Florence-2Production
Microsoft's unified vision foundation model for multiple tasks
- MolmoExperimental
Allen AI's fully open VLM with pointing and grounding
- CogVLM2Experimental
Zhipu AI's vision-language model with video understanding

