Questions? hello@viberation.devGet supportBlogDocsChangelog
Get started

DeepSeek

The open-weight family everyone else gets benchmarked against.

DeepSeek publishes the weights for nearly everything it ships, which is why its models became the price anchor for the whole market — a Flash model here costs a fraction of a frontier model and gets close on reasoning and code. The V4 line is a sparse mixture-of-experts design with a million-token context. This row is the model family; DeepSeek's own free chat app has its own entry in Chats.

At a glance — DeepSeek V4 Flash 0731

Cost
$0.022 in · $0.32 out per 1M tokensBudget
Memory
1.3M tokens — about 1,950 pages
Reads
Text
Can do
  • Use tools (agents)
  • Structured output
  • Thinking mode
Good at
Coding: better than 75% · Agentic work: better than 76% · Overall: better than 66%
Reliability
99.99% uptime, last 24h
Released
Jul 31, 2026
Advanced details
API model ID
deepseek/deepseek-v4-flash-0731
Max output
944K tokens
Cached input
$0.016 per 1M
Tokenizer
DeepSeek
Coding index
69.1
Agentic work index
41.0
Overall index
34.3

Providers

The same model, hosted by different companies. OpenRouter picks one per request and falls back to the next if it fails.

ProviderInput /1MOutput /1MCached /1MMax outputUptime 24h
Relace · fp4$0.022$0.32$0.016944K99.82%
Sail Research · us$0.023$0.42$0.012944K98.86%
Sail Research · fp4$0.023$0.42$0.012944K98.59%
OpenInference · fp8$0.03$0.40$0.008944K98.81%
Inceptron · fp4$0.04$0.271$0.03944K99.94%
StreamLake · fp8$0.044$0.132$0.001384K99.71%
Baidu · fp8$0.06$0.18$0.002131K99.95%
DeepInfra · fp8$0.06$0.18$0.015384K99.89%
Wafer · fast$0.07$0.35$0.06944K99.98%
Reka$0.088$0.528$0.006131K99.72%
Makora$0.09$0.195$0.02384K99.43%
DigitalOcean$0.119$0.238$0.024944K99.85%
BaseTen · fp8$0.13$0.26$0.028384K99.64%
BaseTen · fp8$0.13$0.26$0.028384K99.65%
CoreWeave · fp8$0.13$0.28$0.07236K99.99%
Nebius · fp8$0.14$0.28—922K94.28%
Cohere$0.14$0.28$0.07384K99.01%
Together$0.14$0.28$0.03944K99.61%
Parasail · fp8$0.14$0.28$0.05944K99.79%
Morph$0.142$0.40$0.036944K99.89%
Venice$0.175$0.35$0.03533K99.82%
Mancer 2 · fp8$0.20$0.60—944K99.51%
Fireworks$0.22$0.66$0.007944K98.27%
SiliconFlow · fp8$0.22$0.66$0.028393K99.68%
GMICloud · fp8$0.286$0.858$0.009944K99.51%
Phala$0.308$0.924$0.02393K99.39%
NextBit · fp8$0.352$1.056$0.012944K99.99%
Novita · fp8$0.409$1.228$0.026393K99.99%
AtlasCloud · fp4$0.44$1.32$0.028393K97.87%
Cloudflare$0.44$1.32$0.0141.2M99.99%

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • parallel_tool_calls
  • presence_penalty
  • reasoning
  • reasoning_effort
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_a
  • top_k
  • top_logprobs
  • top_p

Live from OpenRouter, refreshed hourly. Scores from Artificial Analysis.

Key info

Pricing
Open source
Category
Models
Best for
Intermediate
Made by
DeepSeek
Weights
Open
Good for
Reasoning and code at a fraction of frontier prices

More in Models

  • France's frontier lab, with the small models everyone self-hosts.

    Models
  • Meta's current line, built for long multi-agent runs.

    Models
  • Google's open models — Gemini's research, small enough to self-host.

    Models
  • Z.ai's open-weight family, behind the cheap coding subscriptions.

    Models