by Nvidia
Llama-3.1-Nemotron-Ultra-253B is a 253 billion parameter reasoning-focused language model optimized for efficiency that excels at math, coding, and general instruction-following tasks while running on a single 8xH100 node.
Llama-3.1-Nemotron-Ultra-253B is a 253 billion parameter reasoning-focused language model optimized for efficiency that excels at math, coding, and general instruction-following tasks while running on a single 8xH100 node.
On AIHubMix, nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 costs $0.5 per million input tokens and $0.5 per million output tokens.
nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 — no other code changes needed.
nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 is developed by Nvidia. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
NVIDIA-Nemotron-Nano-9B-v2-free is a large language model trained from scratch by NVIDIA…
Developed by Nvidia, Nemotron-Nano-12B-V2-VL-Free is a 12-billion-parameter open…
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model built on a hybrid…
Developed by Nvidia, NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model…
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model featuring…
Developed by NVIDIA, Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal…
Use nvidia/Llama-3_1-Nemotron-Ultra-253B-v1 via the AIHubMix unified API — one interface for every major LLM.