Questions? hello@viberation.devGet supportBlogDocsChangelog
Get started

Nemotron

NVIDIA's open models, tuned for throughput rather than benchmarks.

Nemotron is NVIDIA's open-weight family, built to run fast on NVIDIA hardware. The Lightning models activate only a few billion of their parameters per token and are meant for high-throughput agent workloads, and the line includes specialised members like a content-safety classifier rather than one general model stretched over every job. Several are free to call through OpenRouter.

At a glance — Nemotron 3.5 Lightning

Cost
$0.07 in · $0.20 out per 1M tokensBudgetFree version too
Memory
1M tokens — about 1,500 pages
Reads
Text
Can do
  • Use tools (agents)
  • Structured output
  • Thinking mode
Good at
Coding: better than 28% · Agentic work: better than 28% · Overall: better than 23%
Reliability
99.98% uptime, last 24h
Released
Aug 11, 2026
Advanced details
API model ID
nvidia/nemotron-3.5-lightning
Free version
nvidia/nemotron-3.5-lightning:free
Max output
236K tokens
Cached input
$0.04 per 1M
Tokenizer
Other
Coding index
26.8
Agentic work index
3.5
Overall index
12.9

Providers

The same model, hosted by different companies. OpenRouter picks one per request and falls back to the next if it fails.

ProviderInput /1MOutput /1MCached /1MMax outputUptime 24h
Darkbloom · int4$0.065$0.18—33K99.98%
Io Net$0.07$0.20$0.035131K99.90%
Phala$0.07$0.20$0.04236K99.86%
CoreWeave · bf16$0.07$0.20$0.04236K99.96%
DeepInfra · bf16$0.08$0.20$0.04131K99.75%

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Live from OpenRouter, refreshed hourly. Scores from Artificial Analysis.

Key info

Pricing
Open source
Category
Models
Best for
Expert
Made by
NVIDIA
Weights
Open
Good for
Throughput, and specialised jobs like safety filtering

More in Models

  • France's frontier lab, with the small models everyone self-hosts.

    Models
  • Meta's current line, built for long multi-agent runs.

    Models
  • The open-weight family everyone else gets benchmarked against.

    Models
  • Google's open models — Gemini's research, small enough to self-host.

    Models