Cheapest models for audio
Ranked by input $/Mtoken, then output. Speech and audio models.
- 1Phi-4-multimodal-instruct
Microsoft · 128,000 context
Input
$0/M
Output
$0/M
- 2Llama-3.2-11B-Vision-Instruct
Meta · 128,000 context
Input
$0/M
Output
$0/M
- 3Llama-3.2-90B-Vision-Instruct
Meta · 128,000 context
Input
$0/M
Output
$0/M
- 4magpie-tts-zeroshot
NVIDIA · 1 context
Input
$0/M
Output
$0/M
- 5Nemotron 3 Nano Omni (free)
NVIDIA · 256,000 context
Input
$0/M
Output
$0/M
- 6nemotron-voicechat
NVIDIA · 128,000 context
Input
$0/M
Output
$0/M
- 7Whisper
OpenAI · 1 context
Input
$0/M
Output
$0/M
- 8Whisper Large v3
OpenAI · 1 context
Input
$0/M
Output
$0/M
- 9Auto Router (Beta)
Openrouter · 2,000,000 context
Input
$0/M
Output
$0/M
- 10Auto Router
Openrouter · 2,000,000 context
Input
$0/M
Output
$0/M
- 11Qwen3.5 122B A10B NVFP4
Qwen · 256,144 context
Input
$0/M
Output
$0/M
- 12Qwen3.6 27B FP8
Qwen · 262,144 context
Input
$0/M
Output
$0/M
- 13X-Ai/Grok 4.1 Fast Reasoning
xAI · 20,000,000 context
Input
$0/M
Output
$0/M
- 14X-Ai/Grok-4-Fast-Non-Reasoning
xAI · 2,000,000 context
Input
$0/M
Output
$0/M
- 15X-Ai/Grok-4-Fast-Reasoning
xAI · 2,000,000 context
Input
$0/M
Output
$0/M
- 16Grok STT
xAI · 1 context
Input
$0/M
Output
$0/M
- 17Grok Voice Think Fast 1.0
xAI · 1 context
Input
$0/M
Output
$0/M
- 18KB Whisper
Kblab · 448 context
Input
$0.0023/M
Output
$0.0023/M
- 19Whisper Large v3 Turbo
OpenAI · 448 context
Input
$0.0023/M
Output
$0.0023/M
- 20Gemini-2.0-Flash-Lite
Google · 990,000 context
Input
$0.052/M
Output
$0.210/M
- 21Google: Gemini 2.0 Flash Lite
Google · 1,048,576 context
Input
$0.075/M
Output
$0.300/M
- 22Gemma 4 E2B IT
Google · 32,768 context
Input
$0.100/M
Output
$0.100/M
- 23Voxtral Small 24B 2507
Mistral · 32,000 context
Input
$0.100/M
Output
$0.300/M
- 24Gemini 2.5 Flash-Lite
Gemini · 1,048,576 context
Input
$0.100/M
Output
$0.400/M
- 25Google: Gemini 2.0 Flash
Google · 1,048,576 context
Input
$0.100/M
Output
$0.400/M
- 26Google: Gemini 2.5 Flash Lite Preview 09-2025
Google · 1,048,576 context
Input
$0.100/M
Output
$0.400/M
- 27Gemini Flash-Lite Latest
Google · 1,048,576 context
Input
$0.100/M
Output
$0.400/M
- 28Gemini-2.0-Flash
Google · 990,000 context
Input
$0.100/M
Output
$0.420/M
- 29MiMo-V2.5
Xiaomi · 1,048,576 context
Input
$0.140/M
Output
$0.280/M
- 30Gemma 4 E4B IT
Google · 32,768 context
Input
$0.200/M
Output
$0.200/M
- 31Nemotron 3 Nano Omni 30B A3B Reasoning
NVIDIA · 262,144 context
Input
$0.200/M
Output
$0.800/M
- 32Qwen3 Omni 30B A3B Instruct
Qwen · 65,536 context
Input
$0.250/M
Output
$0.970/M
- 33Qwen3 Omni 30B A3B Thinking
Qwen · 65,536 context
Input
$0.250/M
Output
$0.970/M
- 34Gemini 3.1 Flash Lite Preview
Google · 1,048,576 context
Input
$0.250/M
Output
$1.50/M
- 35Gemini 3.1 Flash Lite
Google · 1,048,576 context
Input
$0.250/M
Output
$1.50/M
- 36Gemini 2.5 Flash
Gemini · 1,048,576 context
Input
$0.300/M
Output
$2.50/M
- 37MiMo-V2.5
Xiaomimimo · 262,144 context
Input
$0.400/M
Output
$2.00/M
- 38MiMo V2 Omni
Xiaomi · 265,000 context
Input
$0.400/M
Output
$2.00/M
- 39Gemini 3 Flash Preview
Google · 1,048,576 context
Input
$0.500/M
Output
$3.00/M
- 40Gemini 3 Flash Preview
Gemini · 1,048,576 context
Input
$0.500/M
Output
$3.00/M