Cohere Transcribe (03-2026)

Released
2026-03-01
Last refreshed
2026-07-01
Status
Researched 91d ago
Open sourceCommercial use: permittedMultimodalVision

Cohere Transcribe (03-2026) is a released vision model with open-source; evaluate it while provider pricing coverage matures.

Cohere releases · 5 in the last 12 months · this family litChangelog →
Specifications
Released
2026-03-01
Parameters
2B
Architecture
Conformer
Specialization
speech-recognition
Openness
Open source
License
Apache 2.0OSI-approvedCommercial use: permitted
Weights
Available
Code
Unknown
Created by

Empowering developers with advanced language AI.

Toronto, Ontario, Canada
Founded 2022
Website
Pricing

No tracked provider token pricing is available yet.

About

Cohere's state-of-the-art automatic speech recognition (ASR) model. Transcribe is a 2B parameter Conformer-based encoder-decoder model trained from scratch for high-fidelity transcription across 14 languages: English, French, German, Italian, Spanish, Portuguese, Greek, Dutch, Polish, Chinese (Mandarin), Japanese, Korean, Vietnamese, and Arabic. Can process 525 minutes of audio per minute. Achieves 5.42 WER on Hugging Face Open ASR leaderboard.

Provider price ladder

No tracked provider token pricing is available for this model yet.

Capabilities

MultimodalAudio

Benchmark peer barsfor Vision

No task-mapped benchmark peers are available for this model yet.

Benchmark scores(2)

Scores are benchmark-specific and are direction-aware: the same numeric gap can mean very different outcomes across suites. Use the leaderboard context and this model's provider route to decide whether the winning margin is meaningful for your workload.
BenchmarkScoreVersionEvaluationSource
LibriSpeech WER (test-clean)1.3test-cleanObserved 2026-03-07—Source
Open ASR Leaderboard (average WER)5.4avg-11-datasetsObserved 2026-03-07—Source

Migration checks

No linked migration route is available for this model yet.