N
NVIDIA / nemotron-3-ultra

NVIDIA Nemotron 3 Ultra

Open 550B-parameter MoE model for orchestration, research, code, and long-duration agents.

01 / SUMMARY

The essential

ORIENTATIONOpen long-duration agents

Primary use stated or inferred cautiously from official documentation.

CONTEXT1,000,000 tokens

Maximum input capacity when published by the source.

EXITNot published

Documented maximum generation limit.

ACCESSNVIDIA NIM and weights

Verified availability channels.

MODEL TYPEText and reasoning · Code

Primary functional family and verified specialties.

LICENSENVIDIA open license

Declared terms for API access, model weights, or self-hosting.

02 / TECHNICAL SHEET

Limits and integration

API IDnemotron-3-ultra
Model typeText and reasoning · Code
Access modelAPI and open weights
LicenseNVIDIA open license
DeploymentNIM, local, or private cloud
Release2026
Knowledge cutoffNot published
EntranceText
ExitText
Context window1,000,000 tokens
maximum outputNot published
ReasoningAgentic reasoning
Published toolsFeatures · NVIDIA NIM
Structured OutingsNot applicable or not published
Batch processingNot published
Prompt cacheNot published
Fine-tuningPublished weights, data, and recipes
Verified platformsNVIDIA NIM · Hugging Face
BEST SUITED FOR

Enterprise agents, code, and accelerated deployment under your own control.

WORTH MONITORING

Its size requires multiple GPUs; compare total cost per task, not just speed per token.

03 / CAPABILITIES

What it can do

01

Orientation

Enterprise agents, code, and accelerated deployment under your own control.

02

Context

1.000.000 tokens · Not published

03

Tools and integration

Functions · NVIDIA NIM

04

Access

NVIDIA NIM and weights

Input modalities
Text
04 / PRICES

Documented cost

ConceptWorthUnit/condition
Standard inputNot publishedStandard API rate
Cached inputNot publishedReading reused prefixes
Cache write or storageNot publishedThe condition varies by provider
Standard outputNot publishedMay include reasoning tokens
Batch inputNot publishedAsynchronous processing
Batch outputNot publishedAsynchronous processing

Consult the primary source: the billing unit depends on the model type and access channel.

i Prices change and may depend on level, region or context length. Check the source before making a decision.

05 / EVALUATIONS

How to read the results

“

A public benchmark provides guidance, but does not replace an evaluation with your data, tools, budget, and error tolerance.

BenchmarkResultMetricSource
Coste en SWE-benchUp to −30 %Published cost per completed taskView source ↗

Inferama only highlights a “best result” when the metric, test set, configuration, and date allow for an equivalent comparison. The supplier's figures are presented as claims from its own source.

06 / VERSIONS

Chronology

NVIDIA Nemotron 3 Ultra

Version added to Inferama's verified catalog.

ANALYSIS

Related analysis

07 / SOURCES

Traceability