Enterprise DNA Enterprise DNA
Directories / Models / Best Audio

Audio

Speech and audio models.

63 models in this board

Category
License
Popular tags

Showing 63 of 63 entries · curated first; switch to All for the indexed long tail

Vision Indexed
Google logo

Gemini 2.5 Flash-Lite

Gemini

Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents

$0.100 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 2.5 Flash

Gemini

Fast Gemini workhorse for multimodal apps where latency and price matter

$0.300 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 2.5 Pro

Gemini

Google's proven reasoning model for coding, math, and multimodal analysis

$1.25 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 3.1 Pro Preview

Gemini

Reasoning-first Gemini preview for agentic coding and complex problem solving

$2.00 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 3.5 Flash

Gemini

Fast Gemini model balancing multimodal reasoning, tool use, and cost

$1.50 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 3 Flash Preview

Gemini

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs

$0.500 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Google: Gemini 2.0 Flash

Google

Fast Gemini model balancing multimodal reasoning, tool use, and cost

$0.100 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Google: Gemini 2.0 Flash Lite

Google

Low-latency Gemini model for high-volume multimodal and agent workloads

$0.075 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini-2.0-Flash-Lite

Google

Low-latency Gemini model for high-volume multimodal and agent workloads

$0.052 / Mtok in 990,000 ctx closed
Vision Indexed
Google logo

Gemini-2.0-Flash

Google

Fast Gemini model balancing multimodal reasoning, tool use, and cost

$0.100 / Mtok in 990,000 ctx closed
Vision Indexed
Google logo

Google: Gemini 2.5 Flash Lite Preview 09-2025

Google

Low-latency Gemini model for high-volume multimodal and agent workloads

$0.100 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 2.5 Pro Preview 05-06

Google

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to

$1.25 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 2.5 Pro Preview 06-05

Google

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to

$1.25 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 3.1 Flash Lite Preview

Google

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2

$0.250 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 3.1 Flash Lite

Google

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and i

$0.250 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 3.1 Pro Preview Custom Tools

Google

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-part

$2.00 / Mtok in 1,048,756 ctx closed
Vision Indexed
Google logo

Gemini 3.1 Pro Preview

Google

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage acro

$2.00 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 3.5 Flash

Google

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficie

$1.50 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 3 Flash Preview

Google

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and t

$0.500 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini 3 Pro Preview

Google

Preview Gemini flagship for complex reasoning, coding, and rich multimodal prompts

$1.25 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Google Gemini Flash Latest

Google

This model always redirects to the latest model in the Google Gemini Flash family.

$1.50 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemini Flash-Lite Latest

Google

Low-latency Gemini model for high-volume multimodal and agent workloads

$0.100 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Google Gemini Pro Latest

Google

This model always redirects to the latest model in the Google Gemini Pro family.

$2.00 / Mtok in 1,048,576 ctx closed
Vision Indexed
Google logo

Gemma 4 E2B IT

Google

Open Gemma instruction model for efficient chat and self-hosted deployments

$0.100 / Mtok in 32,768 ctx open-weights
Vision Indexed
Google logo

Gemma 4 E4B IT

Google

Open Gemma instruction model for efficient chat and self-hosted deployments

$0.200 / Mtok in 32,768 ctx open-weights
Audio Indexed
OpenAI logo

KB Whisper

Kblab

Speech transcription model for accurate audio-to-text and captioning workflows

$0.0023 / Mtok in 448 ctx open-weights
Vision Indexed
Meta logo

Llama-3.2-11B-Vision-Instruct

Meta

Open Llama multimodal model for image understanding and text reasoning

$0 / Mtok in 128,000 ctx open-weights
Vision Indexed
Meta logo

Llama-3.2-90B-Vision-Instruct

Meta

Open Llama multimodal model for image understanding and text reasoning

$0 / Mtok in 128,000 ctx open-weights
Vision Indexed
Meta logo

Muse Spark 1.1

Meta

Open Llama instruction model for multilingual chat, reasoning, and coding

$1.25 / Mtok in 1,048,576 ctx closed
Vision Indexed
Microsoft logo

Phi-4-multimodal-instruct

Microsoft

Multimodal reasoning model for visual analysis, planning, and tool use

$0 / Mtok in 128,000 ctx open-weights
Audio Indexed
Mistral logo

Voxtral Small 24B 2507

Mistral

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech tran

$0.100 / Mtok in 32,000 ctx open-weights
Vision Indexed
NVIDIA logo

magpie-tts-zeroshot

NVIDIA

Speech generation model for controllable voice, narration, and audio delivery

$0 / Mtok in 1 ctx open-weights
Vision Indexed
NVIDIA logo

Nemotron 3 Nano Omni (free)

NVIDIA

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, vide

$0 / Mtok in 256,000 ctx open-weights
Vision Indexed
NVIDIA logo

Nemotron 3 Nano Omni 30B A3B Reasoning

NVIDIA

Open Nemotron omni model combining reasoning with text, vision, and audio

$0.200 / Mtok in 262,144 ctx open-weights
Vision Indexed
NVIDIA logo

nemotron-voicechat

NVIDIA

Nemotron multimodal model for visual reasoning and agentic AI workflows

$0 / Mtok in 128,000 ctx open-weights
Audio Indexed
OpenAI logo

OpenAI: GPT-4o Audio

OpenAI

Speech generation model for controllable voice, narration, and audio delivery

$2.50 / Mtok in 128,000 ctx closed
Audio Indexed
OpenAI logo

GPT-4o mini Transcribe

OpenAI

Speech transcription model for accurate audio-to-text and captioning workflows

$1.25 / Mtok in 1 ctx closed
Audio Indexed
OpenAI logo

GPT-4o Transcribe

OpenAI

Speech transcription model for accurate audio-to-text and captioning workflows

$2.50 / Mtok in 1 ctx closed
Audio Indexed
OpenAI logo

GPT Audio Mini

OpenAI

A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.

$0.600 / Mtok in 128,000 ctx closed
Audio Indexed
OpenAI logo

GPT Audio

OpenAI

The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice con

$2.50 / Mtok in 128,000 ctx closed
Audio Indexed
OpenAI logo

GPT-Realtime-1.5

OpenAI

Speech generation model for controllable voice, narration, and audio delivery

$4.00 / Mtok in 1 ctx closed
Audio Indexed
OpenAI logo

gpt-realtime-2.1

OpenAI

Speech generation model for controllable voice, narration, and audio delivery

$4.00 / Mtok in 128,000 ctx closed
Audio Indexed
OpenAI logo

gpt-realtime-2

OpenAI

Speech generation model for controllable voice, narration, and audio delivery

$4.00 / Mtok in 1 ctx closed
Audio Indexed
OpenAI logo

GPT-Realtime mini

OpenAI

Speech generation model for controllable voice, narration, and audio delivery

$0.600 / Mtok in 1 ctx closed
Audio Indexed
OpenAI logo

Whisper

OpenAI

Speech transcription model for accurate audio-to-text and captioning workflows

$0 / Mtok in 1 ctx closed
Audio Indexed
OpenAI logo

Whisper Large v3 Turbo

OpenAI

Speech transcription model for accurate audio-to-text and captioning workflows

$0.0023 / Mtok in 448 ctx open-weights
Audio Indexed
OpenAI logo

Whisper Large v3

OpenAI

Speech transcription model for accurate audio-to-text and captioning workflows

$0 / Mtok in 1 ctx open-weights
Vision Indexed
O

Auto Router (Beta)

Openrouter

Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#task-spend) for that task based on aggregat

$0 / Mtok in 2,000,000 ctx closed
Vision Indexed
O

Auto Router

Openrouter

Your prompt will be processed by a meta-model and routed to one of dozens of models (see below), optimizing for the best possible output. To see which model was used,...

$0 / Mtok in 2,000,000 ctx closed
Vision Indexed
Qwen logo

Qwen3.5 122B A10B NVFP4

Qwen

Qwen vision-language model for visual reasoning, documents, and agent tasks

$0 / Mtok in 256,144 ctx open-weights
Vision Indexed
Qwen logo

Qwen3.6 27B FP8

Qwen

Qwen vision-language model for visual reasoning, documents, and agent tasks

$0 / Mtok in 262,144 ctx open-weights
Vision Indexed
Qwen logo

Qwen3 Omni 30B A3B Instruct

Qwen

Qwen omni model for text, vision, audio, and multimodal agent tasks

$0.250 / Mtok in 65,536 ctx open-weights
Vision Indexed
Qwen logo

Qwen3 Omni 30B A3B Thinking

Qwen

Qwen omni model for text, vision, audio, and multimodal agent tasks

$0.250 / Mtok in 65,536 ctx open-weights
Vision Indexed
T

Inkling

Thinkingmachines

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning

$1.00 / Mtok in 1,048,576 ctx open-weights
Vision Indexed
xAI logo

X-Ai/Grok 4.1 Fast Reasoning

xAI

Fast Grok model for responsive chat, reasoning, and tool-assisted work

$0 / Mtok in 20,000,000 ctx closed
Vision Indexed
xAI logo

X-Ai/Grok-4-Fast-Non-Reasoning

xAI

Fast Grok model for responsive chat, reasoning, and tool-assisted work

$0 / Mtok in 2,000,000 ctx closed
Vision Indexed
xAI logo

X-Ai/Grok-4-Fast-Reasoning

xAI

Fast Grok model for responsive chat, reasoning, and tool-assisted work

$0 / Mtok in 2,000,000 ctx closed
Audio Indexed
xAI logo

Grok STT

xAI

Speech transcription model for accurate audio-to-text and captioning workflows

$0 / Mtok in 1 ctx closed
Audio Indexed
xAI logo

Grok Voice Think Fast 1.0

xAI

Speech generation model for controllable voice, narration, and audio delivery

$0 / Mtok in 1 ctx closed
Vision Indexed
X

MiMo-V2.5

Xiaomi

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perceptio

$0.140 / Mtok in 1,048,576 ctx open-weights
Vision Indexed
X

MiMo V2 Omni

Xiaomi

MiMo omni model for text, image, video, audio, and agents

$0.400 / Mtok in 265,000 ctx closed
Vision Indexed
X

MiMo-V2.5-Pro

Xiaomimimo

Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution

$1.00 / Mtok in 1,048,576 ctx open-weights
Vision Indexed
X

MiMo-V2.5

Xiaomimimo

Open MiMo model for multimodal coding agents and long-context automation

$0.400 / Mtok in 262,144 ctx open-weights