Voxtral
Mistral's open speech model — audio in, text out.
Voxtral is Mistral's open-weight audio model: it takes speech and files as input and returns text, handling transcription, translation and questions about what was said, while keeping the text ability of the Mistral Small model it is built on. It is a single specialised model rather than a family, so reach for it when your app has recordings to understand, and use a general model for everything else.
At a glance — Voxtral Small 24B 2507
- Cost
- $0.10 in · $0.30 out per 1M tokensBudget
- Memory
- 33K tokens — about 49 pages
- Reads
- Text · Files · Audio
- Can do
- Use tools (agents)
- Structured output
- Thinking mode (not supported)
- Reliability
- 99.49% uptime, last 24h
- Released
- Oct 30, 2025
Advanced details
- API model ID
mistralai/voxtral-small-24b-2507- Max output
- 26K tokens
- Cached input
- $0.01 per 1M
- Tokenizer
- Mistral
Providers
The same model, hosted by different companies. OpenRouter picks one per request and falls back to the next if it fails.
| Provider | Input /1M | Output /1M | Cached /1M | Max output | Uptime 24h |
|---|---|---|---|---|---|
| Mistral · zdr | $0.10 | $0.30 | $0.01 | 26K | 99.42% |
| Mistral | $0.10 | $0.30 | $0.01 | 26K | 99.49% |
| Mistral · eu | $0.11 | $0.33 | $0.011 | 26K | — |
Supported parameters
- frequency_penalty
- max_tokens
- presence_penalty
- response_format
- seed
- stop
- structured_outputs
- temperature
- tool_choice
- tools
- top_p
Live from OpenRouter, refreshed hourly.
Key info
- Pricing
- Open source
- Category
- Models
- Best for
- Expert
- Made by
- Mistral
- Weights
- Open
- Good for
- Transcribing and understanding speech