Cheapest models for cheap
Ranked by input $/Mtoken, then output. Low $/Mtoken for high volume.
- 1KB Whisper
Kblab · 448 context
Input
$0.0023/M
Output
$0.0023/M
- 2Whisper Large v3 Turbo
OpenAI · 448 context
Input
$0.0023/M
Output
$0.0023/M
- 3BGE Reranker Base
Workers Ai · 128,000 context
Input
$0.0031/M
Output
$0/M
- 4Qwen3 Embedding 0.6B
Qwen · 40,960 context
Input
$0.010/M
Output
$0/M
- 5Qwen 3 Embedding 4B
Qwen · 32,000 context
Input
$0.010/M
Output
$0/M
- 6Qwen3-Embedding-8B
Qwen · 32,768 context
Input
$0.010/M
Output
$0/M
- 7Qwen3 Reranker 0.6B
Qwen · 40,960 context
Input
$0.010/M
Output
$0.010/M
- 8Ling-2.6-flash
Inclusionai · 262,144 context
Input
$0.010/M
Output
$0.030/M
- 9BGE M3
Workers Ai · 128,000 context
Input
$0.012/M
Output
$0/M
- 10Qwen3 Embedding 0.6B
Workers Ai · 128,000 context
Input
$0.012/M
Output
$0/M
- 11IBM Granite 4.0 H Micro
Workers Ai · 128,000 context
Input
$0.017/M
Output
$0.110/M
- 12Granite 4.0 Micro
Ibm Granite · 131,000 context
Input
$0.017/M
Output
$0.112/M
- 13Granite 4.0 H Micro
Cf · 131,000 context
Input
$0.017/M
Output
$0.112/M
- 14PLaMo Embedding 1B
Workers Ai · 128,000 context
Input
$0.019/M
Output
$0/M
- 15Mistral Nemo
Mistral · 131,072 context
Input
$0.019/M
Output
$0.030/M
- 16BGE Small EN v1.5
Workers Ai · 128,000 context
Input
$0.020/M
Output
$0/M
- 17E5 Mistral 7B
Intfloat · 4,096 context
Input
$0.020/M
Output
$0.020/M
- 18PaddleOCR-VL
Paddlepaddle · 16,384 context
Input
$0.020/M
Output
$0.020/M
- 19Llama Guard 3 8B
Meta · 131,072 context
Input
$0.020/M
Output
$0.060/M
- 20Sarvam 30B
Sarvam · 128,000 context
Input
$0.020/M
Output
$0.100/M
- 21Manta Flash 1.0
Meganova Ai · 16,384 context
Input
$0.020/M
Output
$0.160/M
- 22Manta Mini 1.0
Meganova Ai · 8,192 context
Input
$0.020/M
Output
$0.160/M
- 23GLM 4.6V FlashX
Z Ai · 200,000 context
Input
$0.020/M
Output
$0.210/M
- 24Mistral Nemo Instruct 2407 TEE
Unsloth · 131,072 context
Input
$0.025/M
Output
$0.098/M
- 25Nex-N2-Mini
Nex Agi · 262,144 context
Input
$0.025/M
Output
$0.100/M
- 26DistilBERT SST-2 INT8
Workers Ai · 128,000 context
Input
$0.026/M
Output
$0/M
- 27Llama 3.2 1B Instruct
Workers Ai · 128,000 context
Input
$0.027/M
Output
$0.200/M
- 28Llama 3.2 1B Instruct
Cf · 60,000 context
Input
$0.027/M
Output
$0.201/M
- 29Llama 3.2 1B Instruct
Meta · 131,072 context
Input
$0.027/M
Output
$0.201/M
- 30Qwen3 4B
Qwen · 128,000 context
Input
$0.030/M
Output
$0.030/M
- 31deepseek/deepseek-ocr-2
DeepSeek · 8,192 context
Input
$0.030/M
Output
$0.030/M
- 32DeepSeek-OCR
DeepSeek · 8,192 context
Input
$0.030/M
Output
$0.030/M
- 33Llama Prompt Guard 2 22M
Meta · 512 context
Input
$0.030/M
Output
$0.030/M
- 34LiquidAI: LFM2-24B-A2B
Liquid · 32,768 context
Input
$0.030/M
Output
$0.120/M
- 35LFM2 24B A2B
Liquidai · 32,768 context
Input
$0.030/M
Output
$0.120/M
- 36gpt-oss-20b
OpenAI · 131,072 context
Input
$0.030/M
Output
$0.130/M
- 37DeepSeek R1 Distill Llama 70B
Deepseek Ai · 131,072 context
Input
$0.030/M
Output
$0.140/M
- 38Doubao-Seed-2.0-mini
Volcengine · 256,000 context
Input
$0.030/M
Output
$0.280/M
- 39AutoGLM-Phone-9B-Multilingual
Zai Org · 65,536 context
Input
$0.035/M
Output
$0.138/M
- 40Nova Micro
Amazon · 128,000 context
Input
$0.035/M
Output
$0.140/M