Questions? hello@viberation.devGet supportBlogDocsChangelog
Get started

Gemma

Google's open models — Gemini's research, small enough to self-host.

Gemma is Google DeepMind's open-weight family, built from the same research as Gemini but published so you can download and run it. Sizes run from a few billion parameters up to about thirty, they read text and images, and the current line has a 256K context with a switchable thinking mode. The obvious starting point if you want a capable model on your own machine or on cheap hardware.

At a glance — Gemma 3 4B

Cost
$0.05 in · $0.10 out per 1M tokensBudget
Memory
131K tokens — about 200 pages
Reads
Text · Images
Can do
  • Use tools (agents) (not supported)
  • Structured output
  • Thinking mode (not supported)
Good at
Coding: better than 0%
Reliability
100.00% uptime, last 24h
Released
Mar 13, 2025
Advanced details
API model ID
google/gemma-3-4b-it
Max output
16K tokens
Knowledge cutoff
2024-08-31
Tokenizer
Gemini
Coding index
2.7

Providers

The same model, hosted by different companies. OpenRouter picks one per request and falls back to the next if it fails.

ProviderInput /1MOutput /1MCached /1MMax outputUptime 24h
DeepInfra · bf16$0.05$0.10—16K100.00%

Supported parameters

  • frequency_penalty
  • logit_bias
  • max_tokens
  • min_p
  • presence_penalty
  • repetition_penalty
  • response_format
  • seed
  • stop
  • structured_outputs
  • temperature
  • top_k
  • top_p

Live from OpenRouter, refreshed hourly. Scores from Artificial Analysis.

Key info

Pricing
Open source
Category
Models
Best for
Intermediate
Made by
Google DeepMind
Weights
Open
Good for
Running a capable model on your own machine

More in Models

  • France's frontier lab, with the small models everyone self-hosts.

    Models
  • Meta's current line, built for long multi-agent runs.

    Models
  • The open-weight family everyone else gets benchmarked against.

    Models
  • Z.ai's open-weight family, behind the cheap coding subscriptions.

    Models