Questions? hello@viberation.devGet supportBlogDocsChangelog
Get started

Nemotron

NVIDIA's open models, tuned for throughput rather than benchmarks.

Nemotron is NVIDIA's open-weight family, built to run fast on NVIDIA hardware. The Lightning models activate only a few billion of their parameters per token and are meant for high-throughput agent workloads, and the line includes specialised members like a content-safety classifier rather than one general model stretched over every job. Several are free to call through OpenRouter.

At a glance — Nemotron 3.5 Content Safety

Cost
$0.20 in · $0.20 out per 1M tokensBudgetFree version too
Memory
131K tokens — about 200 pages
Reads
Text · Images
Can do
  • Use tools (agents) (not supported)
  • Structured output (not supported)
  • Thinking mode
Reliability
100.00% uptime, last 24h
Released
Jun 4, 2026
Advanced details
API model ID
nvidia/nemotron-3.5-content-safety
Free version
nvidia/nemotron-3.5-content-safety:free
Max output
118K tokens
Tokenizer
Other

Providers

The same model, hosted by different companies. OpenRouter picks one per request and falls back to the next if it fails.

ProviderInput /1MOutput /1MCached /1MMax outputUptime 24h
DeepInfra · bf16$0.20$0.20—118K100.00%

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • repetition_penalty
  • seed
  • stop
  • temperature
  • top_k
  • top_p

Live from OpenRouter, refreshed hourly.

Key info

Pricing
Open source
Category
Models
Best for
Expert
Made by
NVIDIA
Weights
Open
Good for
Throughput, and specialised jobs like safety filtering

More in Models

  • France's frontier lab, with the small models everyone self-hosts.

    Models
  • Meta's current line, built for long multi-agent runs.

    Models
  • The open-weight family everyone else gets benchmarked against.

    Models
  • Google's open models — Gemini's research, small enough to self-host.

    Models