Llama 3.1 Nemotron Ultra 253B v1
Reasoning model derived from Meta Llama-3.1-405B-Instruct, post-trained for reasoning, chat preferences, RAG and tool calling.
Specs
- Context
- 131.1K
- Max output
- 16.4K
- Input price
- —
- Output price
- —
- Released
- 2025-04-07
Context
- Context window
- 131.1K
- Max input
- —
- Max output
- 16.4K
Pricing
- Not published
Modalities
- Input
- Text
- Output
- Text
API
- API types
Reasoning
- Reasoning
- Yes
- Levels
- —
Info
- Status
- Active
- Released
- 2025-04-07
- Knowledge cutoff
- —
Features
- Not published
How to call
1 provider
nvidiamodel = nvidia/llama-3.1-nemotron-ultra-253b-v1
JSON
Standard format
{
"id": "llama-3.1-nemotron-ultra-253b-v1",
"object": "model",
"created": 1743984000,
"owned_by": "nvidia",
"name": "Llama 3.1 Nemotron Ultra 253B v1",
"api": {
"types": [
"chat"
]
},
"limits": {
"context": 131072,
"input": null,
"output": 16384
},
"modalities": {
"input": [
"text"
],
"output": [
"text"
]
},
"reasoning": {
"supported": true,
"efforts": []
},
"pricing": null,
"features": [],
"info": {
"status": "active",
"release_date": "2025-04-07",
"knowledge_cutoff": null,
"description": "Reasoning model derived from Meta Llama-3.1-405B-Instruct, post-trained for reasoning, chat preferences, RAG and tool calling.",
"docs": "https://docs.api.nvidia.com/nim/reference/nvidia-llama-3_1-nemotron-ultra-253b-v1",
"verified_at": "2026-10-11"
}
}Official sources
2