NVIDIA Nemotron 3 Ultra in OpenCode: Benchmarks and a Free Setup Guide

By Rakshit Yadav (@yadavrakshit60)•Aug 2026•13 min read

NVIDIA put a 550B open-weight frontier reasoning model behind a free API key at build.nvidia.com, and opencode is an OpenAI-compatible terminal agent that will happily point at it. That combination is the reason this guide exists.

The pitch: a free 1M-context coding model, in a free terminal agent, in about five minutes. This guide is the full walkthrough: what the model actually is, what the benchmarks say (including where it loses), and every command you need end to end.

TL;DR: the four commands

# 1. install opencode
curl -fsSL https://opencode.ai/install | bash

# 2. get a free key at https://build.nvidia.com  (starts with nvapi-)

# 3. add NVIDIA as a provider (config below), then:
opencode

# 4. inside opencode:
/model nvidia/nvidia/nemotron-3-ultra-550b-a55b

Everything below is the detail behind those four lines.

What Nemotron 3 Ultra actually is

NVIDIA released Nemotron 3 Ultra on June 4, 2026, after announcing it at Computex 2026. It is open-weight and downloadable on Hugging Face, which is the part that matters. This is not a closed API-only frontier model.

SpecValue
Total parameters550B
Active parameters55B (Mixture-of-Experts)
ArchitectureHybrid Mamba-Transformer MoE with LatentMoE routing
Context windowUp to 1,000,000 tokens
QuantizationPretrained in NVFP4 (runs on Ampere, Hopper, Blackwell)
ExtrasMulti-token prediction layers, native speculative decoding, inference-time reasoning budget control
Model IDnvidia/nemotron-3-ultra-550b-a55b
ReleaseJune 4, 2026

The design goal NVIDIA states out loud is long-running agents: the Mamba layers keep long sequences cheap, the Transformer layers keep fact retrieval sharp, and the MoE routing means you pay for 55B of compute, not 550B. The NVIDIA technical blog frames the whole model around agentic cost-per-task rather than raw leaderboard position.

That framing is the tell. This is a model built to run in a loop for an hour without bankrupting you, which is exactly the opencode use case.

Benchmarks, with the losses included

Two things to know before reading any table: benchmark numbers for this model vary by source and by reasoning budget, and NVIDIA's own reported ranges are wider than a marketing page usually admits.

NVIDIA's reported numbers (technical blog):

BenchmarkScore
SWE-Bench Verified65-70.4%
Agent productivity (PinchBench)91%
Instruction following (IFBench)82%
Long context (RULER @ 1M)95%
Long-horizon planning (EnterpriseOps-Gym)33%

Third-party aggregation (BenchLM) puts SWE-bench Verified at 71.9%, LiveCodeBench v6 at 89.0%, MMLU-Pro at 86.8%, and IFBench at 81.7%, but ranks the model #154 of 221 overall, with agentic tool use at #120 of 134 and coding at #108 of 138.

Those two paragraphs are the whole story. Read them together:

  • Where it is genuinely strong: instruction following (top-5 in its category), long-context retrieval, single-shot code generation, and raw throughput. NVIDIA claims 5.9x, 4.8x, and 1.6x higher inference throughput than GLM-5.1-754B, Kimi-K2.6-1T, and Qwen-3.5-397B on 8k-in/64k-out, and ~30% lower cost-to-completion on SWE-Bench and Terminal-Bench 2.0 because it burns fewer tokens per turn.
  • Where it is weak: multi-step agentic tool use and long-horizon planning. A 33% on EnterpriseOps-Gym is not a typo. Long-horizon planning is genuinely hard, and this is not where a 550B open model beats a frontier closed model.

So did it kill Claude Code?

No, and the category error is worth naming, because it changes how you should use this.

Claude Code and Codex are harnesses. Nemotron 3 Ultra is a model. Claude Code is a terminal agent that happens to run Claude; opencode is a terminal agent that runs whatever you point it at. The correct comparison is opencode + Nemotron 3 Ultra vs Claude Code + Claude, and the result is exactly what the benchmarks predict:

JobHonest verdict
Generate a file, a function, a migrationNemotron 3 Ultra is completely fine, and free
Read a 300k-token codebase and answer questionsStrong. The 1M context and RULER score are real
Refactor across 20 files in one unattended loopFrontier models still win; agentic tool use is the weak column
Overnight "make CI green"Use a paid frontier model, or expect to babysit
Learning, side projects, unlimited experimentationNemotron wins on price, which is zero

If your monthly AI bill is the thing stopping you from shipping, this stack removes it. If your bottleneck is a hard multi-file refactor, it does not.

The rest of this guide is the setup, which is the same either way.

Step 1: Install opencode

opencode is an open-source terminal coding agent. Pick one:

curl -fsSL https://opencode.ai/install | bash
npm install -g opencode-ai
brew install anomalyco/tap/opencode

Verify it landed:

opencode --version

If the shell says command not found after the curl install, the binary went to ~/.opencode/bin. Add it to your PATH:

echo 'export PATH="$HOME/.opencode/bin:$PATH"' >> ~/.zshrc && source ~/.zshrc

Step 2: Get the free NVIDIA API key

  1. Go to build.nvidia.com and sign in (a free NVIDIA developer account, no credit card).
  2. Open the nemotron-3-ultra-550b-a55b model page.
  3. Click Get API Key / Build with this NIM.
  4. Copy the key. It starts with nvapi-.

Two things worth knowing:

  • The key works across every model on build.nvidia.com, not just Nemotron. Same key, same base URL, different model ID.
  • The endpoint is OpenAI-compatible: https://integrate.api.nvidia.com/v1. Anything that speaks the OpenAI chat-completions API can use it.

Sanity-check the key before you touch opencode config. This saves you debugging the wrong layer:

curl https://integrate.api.nvidia.com/v1/chat/completions -H "Authorization: Bearer $NVIDIA_API_KEY" -H "Content-Type: application/json" -d '{"model":"nvidia/nemotron-3-ultra-550b-a55b","messages":[{"role":"user","content":"say hi"}],"max_tokens":20}'

If that returns JSON with a message, the key is good. If it returns 401, the key is wrong. If it returns 404, the model ID is wrong.

Step 3: Add NVIDIA as an opencode provider

opencode does not ship an NVIDIA provider by default, so you add it as a generic OpenAI-compatible provider. Two files.

Config, at ~/.config/opencode/opencode.json:

{
  "$schema": "https://opencode.ai/config.json",
  "provider": {
    "nvidia": {
      "npm": "@ai-sdk/openai-compatible",
      "name": "NVIDIA NIM",
      "options": {
        "baseURL": "https://integrate.api.nvidia.com/v1"
      },
      "models": {
        "nvidia/nemotron-3-ultra-550b-a55b": {
          "name": "Nemotron 3 Ultra 550B",
          "limit": {
            "context": 1000000,
            "output": 32768
          }
        }
      }
    }
  }
}

Credentials, at ~/.local/share/opencode/auth.json:

{
  "nvidia": {
    "type": "api",
    "key": "nvapi-YOUR_KEY_HERE"
  }
}

Keeping the key in auth.json instead of opencode.json is deliberate: the config file is the one you might commit to a dotfiles repo, and the auth file is the one you never do.

The provider key (nvidia) must be identical in both files. That mismatch is the single most common setup failure.

Optional: reasoning budget

Nemotron 3 Ultra supports inference-time reasoning control. If you want it thinking harder, or cheaper and faster, pass it through extraBody, and raise the timeout, because long reasoning turns will blow past opencode's default:

"models": {
  "nvidia/nemotron-3-ultra-550b-a55b": {
    "name": "Nemotron 3 Ultra 550B",
    "options": {
      "timeout": 600000,
      "extraBody": {
        "chat_template_kwargs": {
          "thinking": true,
          "reasoning_budget": 512
        }
      }
    },
    "limit": { "context": 1000000, "output": 32768 }
  }
}

Lower reasoning_budget for speed on simple edits; raise it for hard debugging. This is the knob that most changes the feel of the model.

Step 4: Select the model and run

cd ~/your-project
opencode

Inside the TUI:

/init

/init reads your repo and writes an AGENTS.md, the project brief the agent reads on every session. Do not skip it. An agent with no project conventions will invent them.

Then pick the model:

/model

Type nemotron to filter, hit enter. Or set it directly:

/model nvidia/nvidia/nemotron-3-ultra-550b-a55b

The doubled nvidia/nvidia/ is not a typo. The first segment is your provider key from opencode.json, the second is part of NVIDIA's model ID.

Want it to be the default so you never pick again? Add to opencode.json:

"model": "nvidia/nvidia/nemotron-3-ultra-550b-a55b"

Now give it real work:

Read AGENTS.md, then add a /health endpoint that returns build SHA and uptime. Follow the existing route conventions.

Troubleshooting

SymptomCauseFix
Provider not foundKey mismatch between opencode.json and auth.jsonBoth must say nvidia
401 UnauthorizedBad or expired keyRegenerate at build.nvidia.com; confirm the nvapi- prefix
404 model not foundWrong model IDIt is nvidia/nemotron-3-ultra-550b-a55b, lowercase
Hangs on "Thinking" foreverLong reasoning turn exceeding the default timeoutSet "timeout": 600000 and lower reasoning_budget
Model not in /model listConfig not reloaded, or invalid JSONFully quit and relaunch opencode; run the file through jq .
429 rate limitedFree tier throttleWait, or fall back to a second provider
Tool calls loop or stall mid-taskThe known weak column: agentic tool useBreak the task into smaller explicit steps

That last row is not a bug you can configure away. It is the benchmark showing up in your terminal. Smaller, more explicit prompts are the workaround.

Free tier reality check

NVIDIA's NIM tier is free and requires no credit card. What NVIDIA does not clearly publish is the exact RPM/RPD limits or firm commercial terms, so treat it as a generous developer tier, not a production SLA. Check the current terms before you ship a product on it.

If you want a documented number for comparison, OpenRouter's free tier for the same model publishes 20 requests/minute and 200 requests/day. That is a reasonable mental model for what "free tier" means here: fine for a solo developer working a normal day, not fine for a CI pipeline hammering it.

Since the weights are open, the third option always exists: self-host. Realistically that means 8x H100 or 4x B200, so for most people, the free API is the answer.

What I would actually do with this

A stack that costs nothing changes what you're willing to try. The specific wins:

  1. Learning and side projects. Unlimited experimentation with zero billing anxiety is the entire point.
  2. Large-codebase Q&A. The 1M context and the RULER score are the strongest real claim here. Point it at a repo you did not write and ask it questions.
  3. A second opinion. Run the same bug past a different model family. Disagreement between two models is a genuinely useful signal.
  4. Bulk mechanical work. Test backfills, docstrings, type annotations: high volume, low ambiguity, exactly the shape it handles well.

Where I would still reach for a paid frontier model: unattended multi-file refactors and anything where a wrong tool call costs an hour.

And the part that does not change with the model: the agent is only as good as the context you hand it. AGENTS.md, scoped rules, and a review step matter more than which weights are behind the API. That is the same argument as Cursor vs Claude Code: pick by job, and keep the project's brain in files on disk.

Sources

Related Tools & Agents

🛠️ Free Tool: kit-picker🤖 Agent: tech-lead🤖 Agent: code-reviewer⚡ Command: /ship⚡ Command: /fixSkill: project-conventionsSkill: prompt-engineering

A free model still needs a project brain

Agents follow conventions you write down. See how the kit installs rules, agents, and commands into a repo.

Read the install docs