by OpenAI
GPT-4o (“o” stands for “omni”) is a new-generation multimodal model designed for more natural human–computer interaction. It can accept any combination of text, audio, image, and video as input, and generate multimodal outputs including text, audio, and images. With audio response latency as low as 232 milliseconds on average around 320 milliseconds, it approaches real human conversational speed. The model delivers strong performance in English text and code, significantly improved multilingual understanding, and outstanding capabilities in visual and audio perception, while offering faster API performance and substantially reduced cost for real-time and complex multimodal applications.
GPT-4o (“o” stands for “omni”) is a new-generation multimodal model designed for more natural human–computer interaction. It can accept any combination of text, audio, image, and video as input, and generate multimodal outputs including text, audio, and images. With audio response latency as low as 232 milliseconds on average around 320 milliseconds, it approaches real human conversational speed. The model delivers strong performance in English text and code, significantly improved multilingual understanding, and outstanding capabilities in visual and audio perception, while offering faster API performance and substantially reduced cost for real-time and complex multimodal applications.
gpt-4o has a 128,000 token context window.
On AIHubMix, gpt-4o costs $2.5 per million input tokens and $10 per million output tokens. Cached input reads are billed at $1.25 per million tokens.
gpt-4o accepts text and image input.
gpt-4o supports tool calling, function calling and structured outputs. Per-protocol parameter support is listed in the capability table on this page.
gpt-4o is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to gpt-4o — no other code changes needed.
gpt-4o is developed by OpenAI. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Free version: gpt-4o-free
GPT-5.6 Luna is designed for cost-sensitive, high-volume workloads. It roughly…
GPT‑5.6 Sol sets a new standard for both intelligence and efficiency, achieving…
GPT-5.6 Terra is designed for workloads that balance intelligence and cost. It roughly…
GPT-4o Transcribe Diarize is an automatic speech recognition (ASR) model with built-in…
The gpt-audio model is OpenAI's first officially released (generally available) audio…
GPT-image-2 is OpenAI's latest cutting-edge image generation model. Key value adds…
Use gpt-4o via the AIHubMix unified API — one interface for every major LLM.