OpenAI's multimodal flagship — text, vision, audio in a single model with fast inference and broad capability.
Five passes over the same idea, each from a different angle. Do them in order, or jump to whichever you need.
GPT-4o is OpenAI's omni model supporting text, image, and audio inputs with a single architecture. It offers competitive pricing, 128k context, native tool calling, JSON mode, and strong performance across coding, reasoning, and creative tasks. Understanding its strengths and limitations helps you choose the right model for each use case.