LLM Reference
AA ASR WERactiveAudio

AA ASR WER: Artificial Analysis ASR WER

Metric: WER (%) (lower is better)Introduced: 2024

Word Error Rate measured by Artificial Analysis using a consistent methodology across a proprietary multi-domain test suite. Provides independent, reproducible comparison across both open-source and commercial STT APIs. Lower is better.

Models ranked

4

tracked on this benchmark

Score band

2.4 – 3.6

best → lowest tracked

Snapshot trend

-0.00

Apr 2 → Aug 26 · 1 models

Leaderboard

Tracked models ranked by WER (%) (lower is better).

Compare candidates
#Model variant and provenanceScore
1
MAI-Transcribe-1.5
Version: aa-werHarness: Not recordedEvaluator: Not recordedObserved: Jun 2, 2026Confidence: Not recordedSource
2.4
2
MAI-Transcribe-1
Version: aa-werHarness: Not recordedEvaluator: Not recordedObserved: Apr 2, 2026Confidence: Not recordedSource
2.6
3
Gemini 3.5 Transcribe
Version: AA-WER; non-streaming/batch; result as of August 2026Harness: Not recordedEvaluator: Not recordedObserved: Aug 26, 2026Confidence: confirmedSource

Notes: Recommended seed value: 2.6 WER (%; lower is better). Evaluator: Artificial Analysis. Harness/methodology: Google reports locally evaluating default competitor APIs; Artificial Analysis' AA-WER index spans diverse datasets and separately evaluates non-streaming and streaming modalities. This row is for gemini-3.5-transcribe only, not the Live endpoint.

2.6
4
Voxtral Mini Transcribe 2
Version: aa-werHarness: Not recordedEvaluator: Not recordedObserved: Feb 5, 2026Confidence: Not recordedSource
3.6

How to read this benchmark

This benchmark scores models where lower is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

Related benchmarks

Last reviewed: Jun 7, 2026

Resources