by OpenAI
Ultra-lightweight model with million-token context, optimized for speed and low latency, costing only $0.10 per million input tokens. It is suitable for edge computing and real-time interaction. The automatic caching mechanism offers a 75% cost reduction on cache hits.
Ultra-lightweight model with million-token context, optimized for speed and low latency, costing only $0.10 per million input tokens. It is suitable for edge computing and real-time interaction. The automatic caching mechanism offers a 75% cost reduction on cache hits.
gpt-4.1-nano has a 1,047,576 token context window.
On AIHubMix, gpt-4.1-nano costs $0.1 per million input tokens and $0.4 per million output tokens. Cached input reads are billed at $0.03 per million tokens.
gpt-4.1-nano accepts text and image input.
gpt-4.1-nano supports tool calling, function calling, structured outputs and long context. Per-protocol parameter support is listed in the capability table on this page.
gpt-4.1-nano is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to gpt-4.1-nano — no other code changes needed.
gpt-4.1-nano is developed by OpenAI. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Free version: gpt-4.1-nano-free
GPT-5.6 Luna is designed for cost-sensitive, high-volume workloads. It roughly…
GPT‑5.6 Sol sets a new standard for both intelligence and efficiency, achieving…
GPT-5.6 Terra is designed for workloads that balance intelligence and cost. It roughly…
GPT-4o Transcribe Diarize is an automatic speech recognition (ASR) model with built-in…
The gpt-audio model is OpenAI's first officially released (generally available) audio…
GPT-image-2 is OpenAI's latest cutting-edge image generation model. Key value adds…
Use gpt-4.1-nano via the AIHubMix unified API — one interface for every major LLM.