Cheapest models for vision
Ranked by input $/Mtoken, then output. Image and multimodal understanding.
- 1Claude Sonnet 5 (Free)
Anthropic · 1,000,000 context
Input
$0/M
Output
$0/M
- 2Kimi K2.7 Code (Free)
Moonshotai · 262,144 context
Input
$0/M
Output
$0/M
- 3glm-4.7
Novita · 205,000 context
Input
$0/M
Output
$0/M
- 4kimi-k2-thinking
Novita · 256,000 context
Input
$0/M
Output
$0/M
- 5minimax-m2.1
Novita · 205,000 context
Input
$0/M
Output
$0/M
- 6Step 3.7 Flash (Free)
Stepfun · 256,000 context
Input
$0/M
Output
$0/M
- 7Gemma 4 31B (free)
Google · 262,144 context
Input
$0/M
Output
$0/M
- 8GLM-4.6
Novita · 1 context
Input
$0/M
Output
$0/M
- 9Gemma 4 26B A4B (free)
Google · 262,144 context
Input
$0/M
Output
$0/M
- 10glm-4.7-flash
Novita · 200,000 context
Input
$0/M
Output
$0/M
- 11glm-4.6v
Novita · 131,000 context
Input
$0/M
Output
$0/M
- 12Mistral Small 3.1
Mistral Ai · 128,000 context
Input
$0/M
Output
$0/M
- 13Mistral Medium 3 (25.05)
Mistral Ai · 128,000 context
Input
$0/M
Output
$0/M
- 14Phi-4-multimodal-instruct
Microsoft · 128,000 context
Input
$0/M
Output
$0/M
- 15Mistral Large 3 675B Instruct 2512
Mistral · 262,144 context
Input
$0/M
Output
$0/M
- 16FLUX.1-Kontext-dev
Black Forest Labs · 40,960 context
Input
$0/M
Output
$0/M
- 17Seedance 2.0 Fast
Bytedance · 1 context
Input
$0/M
Output
$0/M
- 18Seedance 2.0
Bytedance · 1 context
Input
$0/M
Output
$0/M
- 19Seedance 2
Bytedance · 4,096 context
Input
$0/M
Output
$0/M
- 20llama-3.3-70b-cs
Cerebras · 1 context
Input
$0/M
Output
$0/M
- 21qwen3-235b-2507-cs
Cerebras · 1 context
Input
$0/M
Output
$0/M
- 22qwen3-32b-cs
Cerebras · 1 context
Input
$0/M
Output
$0/M
- 23deepseek-ai/DeepSeek-OCR
Deepseek Ai · 8,192 context
Input
$0/M
Output
$0/M
- 24ElevenLabs-Music
Elevenlabs · 2,000 context
Input
$0/M
Output
$0/M
- 25ElevenLabs-v2.5-Turbo
Elevenlabs · 128,000 context
Input
$0/M
Output
$0/M
- 26ElevenLabs-v3
Elevenlabs · 128,000 context
Input
$0/M
Output
$0/M
- 27Kimi-K2.5-FW
Fireworks Ai · 262,144 context
Input
$0/M
Output
$0/M
- 28Gemma 3n E2b It
Google · 128,000 context
Input
$0/M
Output
$0/M
- 29Gemma 4 31B IT FP8
Google · 262,144 context
Input
$0/M
Output
$0/M
- 30paligemma
Google · 128,000 context
Input
$0/M
Output
$0/M
- 31Imagen-3-Fast
Google · 480 context
Input
$0/M
Output
$0/M
- 32Imagen-3
Google · 480 context
Input
$0/M
Output
$0/M
- 33Imagen-4-Fast
Google · 480 context
Input
$0/M
Output
$0/M
- 34Imagen-4-Ultra
Google · 480 context
Input
$0/M
Output
$0/M
- 35Imagen-4
Google · 480 context
Input
$0/M
Output
$0/M
- 36Lyria 3 Clip Preview
Google · 1,048,576 context
Input
$0/M
Output
$0/M
- 37Lyria 3 Pro Preview
Google · 1,048,576 context
Input
$0/M
Output
$0/M
- 38Lyria
Google · 1 context
Input
$0/M
Output
$0/M
- 39Veo-2
Google · 480 context
Input
$0/M
Output
$0/M
- 40Veo-3.1-Fast
Google · 480 context
Input
$0/M
Output
$0/M