Questions? hello@viberation.devGet supportBlogDocsChangelog
Get started

DeepSeek

The open-weight family everyone else gets benchmarked against.

DeepSeek publishes the weights for nearly everything it ships, which is why its models became the price anchor for the whole market — a Flash model here costs a fraction of a frontier model and gets close on reasoning and code. The V4 line is a sparse mixture-of-experts design with a million-token context. This row is the model family; DeepSeek's own free chat app has its own entry in Chats.

At a glance — DeepSeek V4 Flash 0423

Cost
$0.047 in · $0.094 out per 1M tokensBudget
Memory
1M tokens — about 1,550 pages
Reads
Text
Can do
  • Use tools (agents)
  • Structured output
  • Thinking mode
Good at
Coding: better than 65% · Agentic work: better than 56% · Overall: better than 46%
Reliability
99.97% uptime, last 24h
Released
Apr 24, 2026
Advanced details
API model ID
deepseek/deepseek-v4-flash
Max output
384K tokens
Max input
1M tokens
Cached input
$0.009 per 1M
Tokenizer
DeepSeek
Coding index
56.2
Agentic work index
22.2
Overall index
24.2

Providers

The same model, hosted by different companies. OpenRouter picks one per request and falls back to the next if it fails.

ProviderInput /1MOutput /1MCached /1MMax outputUptime 24h
Relace · fp4$0.03$0.50$0.016944K99.54%
OpenInference · fp8Degraded$0.04$0.50$0.014944K96.34%
Baidu · fp8$0.047$0.094$0.009131K97.58%
StreamLake · fp8$0.047$0.094$0.009384K95.81%
DeepInfra · fp8$0.09$0.18$0.01866K99.75%
GMICloud · fp8$0.091$0.182$0.018944K99.47%
Venice$0.097$0.193$0.0233K99.12%
DigitalOcean$0.098$0.196$0.02384K99.85%
SiliconFlow · fp8$0.13$0.28$0.028393K99.50%
Alibaba · fp8$0.134$0.268$0.027393K98.05%
Novita · fp8$0.14$0.28$0.028393K99.97%
AtlasCloud · fp4$0.14$0.28$0.028393K98.79%
Parasail · fp8$0.14$0.28$0.07944K99.77%
NextBit · fp8$0.15$0.30$0.035944K99.84%
Mancer 2 · fp8$0.19$0.50—944K97.75%
Azure · us$0.21$0.56$0.031384K97.96%

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_completion_tokens
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • reasoning_effort
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_a
  • top_k
  • top_logprobs
  • top_p

Live from OpenRouter, refreshed hourly. Scores from Artificial Analysis.

Key info

Pricing
Open source
Category
Models
Best for
Intermediate
Made by
DeepSeek
Weights
Open
Good for
Reasoning and code at a fraction of frontier prices

More in Models

  • France's frontier lab, with the small models everyone self-hosts.

    Models
  • Meta's current line, built for long multi-agent runs.

    Models
  • Google's open models — Gemini's research, small enough to self-host.

    Models
  • Z.ai's open-weight family, behind the cheap coding subscriptions.

    Models