Skip to content
Not available in this workspace
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Support
  • Works With OR
  • Data

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube
Collections/Speech-to-Text

Best Speech-to-Text and Transcription Models

Model rankings updated August 2026 based on real usage data.

Speech-to-text models convert spoken audio into text for transcription, captions, meeting summaries, call analysis, and voice-driven applications. This collection ranks transcription models by their usage on OpenRouter over the past week. The current top models are GPT-4o Mini Transcribe, GPT-4o Transcribe, and Voxtral Mini Transcribe. Compare accuracy, speed, language support, and cost to match the right model to your audio workflow.

Browse All ModelsCompare Models

Top Speech-to-Text Models on OpenRouter

Favicon for openai

OpenAI: GPT-4o Mini Transcribe

220M tokens

GPT-4o Mini Transcribe is OpenAI's smaller, cost-efficient speech-to-text model built on GPT-4o Mini audio capabilities. It's priced per token (input and output), making it suitable for high-volume transcription workflows that benefit from token-level billing transparency at a lower cost point.

by openai128K context$1.25/M input tokens$5/M output tokens
Favicon for openai

OpenAI: GPT-4o Transcribe

86.1M tokens

GPT-4o Transcribe is OpenAI's high-quality speech-to-text model built on GPT-4o audio capabilities. It's priced per token (input and output), making it suitable for workflows that benefit from token-level billing transparency.

by openai128K context$2.50/M input tokens$10/M output tokens
Favicon for mistralai

Mistral: Voxtral Mini Transcribe

11.8M tokens

Voxtral Mini Transcribe is Mistral's speech-to-text model, derived from the Voxtral Mini family. It accepts audio input and returns transcribed text via the standard transcription API. Suited for transcribing meetings, voice notes, podcasts, and other spoken content.

by mistralai$0.003/minute
Favicon for nvidia

NVIDIA: Nemotron 3.5 ASR Streaming Multilingual 0.6B

Nemotron 3.5 ASR Streaming Multilingual 0.6B is a speech recognition model from NVIDIA. Its prompt-conditioned, cache-aware FastConformer-RNNT design targets low-latency transcription across more than 40 languages for real-time captioning, voice agents, and multilingual transcription pipelines.

by nvidia$0.000003/second
Favicon for mistralai

Mistral: Voxtral Small 24B 2507 STT

Voxtral Small 24B 2507 STT is a speech transcription model from Mistral AI. It is suited for transcription, translation, and audio understanding workloads that benefit from its larger model capacity.

by mistralai$0.00005/second
Favicon for mistralai

Mistral: Voxtral Mini 3B 2507

Voxtral Mini 3B 2507 is a speech and audio understanding model from Mistral AI. It is suited for transcription, translation, and compact audio processing workloads.

by mistralai$0.000017/second
Favicon for qwen

Qwen: Qwen3 ASR 1.7B

Qwen3 ASR 1.7B is an automatic speech recognition model from Qwen. It supports multilingual language identification and transcription across 30 languages and 22 Chinese dialects, with streaming and offline inference plus segment-level and word-level timestamps.

by qwen$0.000008/second
Favicon for qwen

Qwen: Qwen3 ASR 0.6B

Qwen3 ASR 0.6B is a compact automatic speech recognition model from Qwen. It supports multilingual language identification and transcription across 30 languages and 22 Chinese dialects, with streaming and offline inference plus segment-level and word-level timestamps.

by qwen$0.000003/second
Favicon for openai

OpenAI: GPT Transcribe

GPT Transcribe is a high-accuracy speech-to-text model from OpenAI. It is suited for recorded audio, streamed file transcription, and committed Realtime turns, with free-form context, keyword hints, and multiple language hints for specialized terms and multilingual speech.

by openai$0.0045/minute
Favicon for fish-audio

Fish Audio: Transcribe 1

Transcribe 1 is a speech-to-text model from Fish Audio. It is suited for audio transcription with automatic language detection and can return timestamped word-level segments when alignment details are requested.

by fish-audio$0.0001/second

Explore more collections

  • Free Models
  • Discounted Models
  • Coding
  • Roleplay
  • Vision Models
  • Tool Calling
  • OpenClaw
  • Image Models
  • Video Models
  • Audio Models
  • Text-to-Speech
  • Embedding Models
  • Rerank Models
  • Distillable Models
  • All collections