Questions? hello@viberation.devGet supportBlogDocsChangelog
Get started

Ling

inclusionAI's open mixture-of-experts models, priced near zero.

Ling is inclusionAI's open-weight family: mixture-of-experts models that activate a small slice of their parameters per token, which is how the Flash models land at a price per million tokens most people would call a rounding error. There are vision and finance-tuned variants alongside the general one, and a 262K context. A sensible pick for high-volume work — classifying, tagging, summarising — where a frontier model is overkill.

At a glance — Ling 3.0 Flash VL

Cost
$0.021 in · $0.062 out per 1M tokensBudget
Memory
262K tokens — about 400 pages
Reads
Text · Images · Video
Can do
  • Use tools (agents)
  • Structured output
  • Thinking mode
Good at
Coding: better than 66% · Agentic work: better than 65% · Overall: better than 46%
Reliability
99.92% uptime, last 24h
Released
Sep 10, 2026
Advanced details
API model ID
inclusionai/ling-3.0-flash-vl
Max output
33K tokens
Cached input
$0.004 per 1M
Tokenizer
Other
Coding index
57.0
Agentic work index
28.7
Overall index
24.6

Providers

The same model, hosted by different companies. OpenRouter picks one per request and falls back to the next if it fails.

ProviderInput /1MOutput /1MCached /1MMax outputUptime 24h
Novita · bf16$0.021$0.062$0.00433K99.92%
DeepInfra · fp16$0.06$0.18$0.01233K99.61%

Supported parameters

  • frequency_penalty
  • include_reasoning
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • presence_penalty
  • reasoning
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Live from OpenRouter, refreshed hourly. Scores from Artificial Analysis.

Key info

Pricing
Open source
Category
Models
Best for
Expert
Made by
inclusionAI
Weights
Open
Good for
High-volume, low-value calls

More in Models

  • France's frontier lab, with the small models everyone self-hosts.

    Models
  • Meta's current line, built for long multi-agent runs.

    Models
  • The open-weight family everyone else gets benchmarked against.

    Models
  • Google's open models — Gemini's research, small enough to self-host.

    Models