Molmo
Allen AI's fully open VLM with pointing and grounding
Verdict
Truly open VLM with open weights, data, and code. Unique pointing capability for spatial grounding. Good for academic research and applications requiring full transparency.
Other Multimodal & Vision
- LLaVA-NeXTStable
Leading open-source vision-language model with strong reasoning
- Pixtral LargeStable
Mistral's 124B vision-language model with 128K context
- InternVL3Stable
Open-source VLM rivalling closed frontier models on vision benchmarks
- Qwen-VL-MaxStable
Alibaba's flagship vision-language model with video understanding
- Florence-2Production
Microsoft's unified vision foundation model for multiple tasks
- CogVLM2Experimental
Zhipu AI's vision-language model with video understanding

