Questions? hello@viberation.devGet supportBlogDocsChangelog
Get started

Llama

Meta's open-weight family, and the reason local models took off.

Llama is Meta's open-weight family and the one that made running a serious model on your own hardware normal — most local tooling was built for it first, so support is everywhere. The Llama 4 models are mixture-of-experts and multimodal, with very large context windows, and the line includes Llama Guard, a classifier for filtering input and output rather than a chat model. Meta's newer work ships under a different name, so check the dates on this family before choosing it.

At a glance — Llama 4 Maverick

Cost
$0.188 in · $0.653 out per 1M tokensBudget
Memory
1M tokens — about 1,550 pages
Reads
Text · Images
Can do
  • Use tools (agents)
  • Structured output
  • Thinking mode (not supported)
Good at
Coding: better than 15% · Agentic work: better than 2%
Reliability
99.86% uptime, last 24h
Released
Apr 5, 2025
Advanced details
API model ID
meta-llama/llama-4-maverick
Max output
16K tokens
Knowledge cutoff
2024-08-31
Tokenizer
Llama4
Coding index
16.3
Agentic work index
0.6

Providers

The same model, hosted by different companies. OpenRouter picks one per request and falls back to the next if it fails.

ProviderInput /1MOutput /1MCached /1MMax outputUptime 24h
DigitalOcean$0.188$0.653—16K99.86%
DeepInfra · base$0.20$0.80—16K99.32%
Novita · fp8$0.27$0.85—8K99.64%
Parasail · fp8$0.35$1.00$0.1733K99.55%
Google · us-east5$0.35$1.15—8K—

Supported parameters

  • frequency_penalty
  • logit_bias
  • logprobs
  • max_tokens
  • min_p
  • presence_penalty
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • tool_choice
  • tools
  • top_k
  • top_logprobs
  • top_p

Live from OpenRouter, refreshed hourly. Scores from Artificial Analysis.

Key info

Pricing
Open source
Category
Models
Best for
Intermediate
Made by
Meta
Weights
Open
Good for
Local setups, where tooling support is widest

More in Models

  • France's frontier lab, with the small models everyone self-hosts.

    Models
  • Meta's current line, built for long multi-agent runs.

    Models
  • The open-weight family everyone else gets benchmarked against.

    Models
  • Google's open models — Gemini's research, small enough to self-host.

    Models