Model comparison
GPT-OSS-120B-CS vs Llama-3.1-8B-CS
Compare GPT-OSS-120B-CS and Llama-3.1-8B-CS: input/output $/Mtoken, context window, modalities, and license. OpenRouter-synced pricing.
Cerebras
GPT-OSS-120B-CS
Open-weight GPT model for self-hosted reasoning and instruction-following workloads
Cerebras
Llama-3.1-8B-CS
Open Llama instruction model for multilingual chat, reasoning, and coding
| Metric | GPT-OSS-120B-CS | Llama-3.1-8B-CS |
|---|---|---|
| Provider | Cerebras | Cerebras |
| Context window | 128,000 | 128,000 |
| Input $/Mtok | $0.350 | $0.100 |
| Output $/Mtok | $0.750 | $0.100 |
| Modalities | text, vision | text, vision |
| License | closed | closed |
Quick take
On input price, Llama-3.1-8B-CS is cheaper at $0.100/Mtok. For context window, GPT-OSS-120B-CS leads with 128,000 tokens.
Pick based on your workload: high-volume cheap inference vs long-document / agent loops. Enterprise DNA can wire either model into Omni with evals, secrets, and job orchestration.