Questions? hello@viberation.devGet supportBlogDocsChangelog
Get started

Nemotron

NVIDIA's open models, tuned for throughput rather than benchmarks.

Nemotron is NVIDIA's open-weight family, built to run fast on NVIDIA hardware. The Lightning models activate only a few billion of their parameters per token and are meant for high-throughput agent workloads, and the line includes specialised members like a content-safety classifier rather than one general model stretched over every job. Several are free to call through OpenRouter.

At a glance — Nemotron 3 Nano 30B A3B

Cost
$0.05 in · $0.20 out per 1M tokensBudget
Memory
262K tokens — about 400 pages
Reads
Text
Can do
  • Use tools (agents)
  • Structured output
  • Thinking mode
Good at
Coding: better than 11% · Agentic work: better than 12% · Overall: better than 6%
Reliability
100.00% uptime, last 24h
Released
Dec 14, 2025
Advanced details
API model ID
nvidia/nemotron-3-nano-30b-a3b
Max output
236K tokens
Cached input
$0.03 per 1M
Tokenizer
Other
Coding index
14.4
Agentic work index
1.0
Overall index
8.9

Providers

The same model, hosted by different companies. OpenRouter picks one per request and falls back to the next if it fails.

ProviderInput /1MOutput /1MCached /1MMax outputUptime 24h
Crusoe · fp8$0.05$0.20$0.03236K100.00%
Novita · fp4$0.05$0.20—33K99.96%
DeepInfra · fp4$0.05$0.20$0.025228K99.79%
Nebius · fp8$0.06$0.24—236K98.01%

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Live from OpenRouter, refreshed hourly. Scores from Artificial Analysis.

Key info

Pricing
Open source
Category
Models
Best for
Expert
Made by
NVIDIA
Weights
Open
Good for
Throughput, and specialised jobs like safety filtering

More in Models

  • France's frontier lab, with the small models everyone self-hosts.

    Models
  • Meta's current line, built for long multi-agent runs.

    Models
  • The open-weight family everyone else gets benchmarked against.

    Models
  • Google's open models — Gemini's research, small enough to self-host.

    Models