How Accurate Is AssemblyAI? Universal-3.5 Pro Benchmarks, Independently Checked
AssemblyAI's current flagship is Universal-3.5 Pro (July 2026), which supersedes the now-deprecated Universal-3 Pro (February 2026 — the AssemblyAI API rejects new U-3 Pro requests). In our own July 2026 first-hand benchmark of 903 audio files across 19 datasets, Universal-3.5 Pro ranked first among 9 tested promptable AI-transcription APIs on aggregate accuracy (7.0% average WER), winning 11 of 15 head-to-head comparisons against Universal-2. Expanded testing that added Speechmatics found its Melia-1 model at 6.4% aggregate WER, edging Universal-3.5 Pro on pure batch accuracy — U-3.5 Pro remains the leading promptable AI-transcription API (Speechmatics does not support prompting or keyterms). Historically, AssemblyAI's previous flagship Universal-3 Pro measured 2.3% WER on Artificial Analysis's AgentTalk benchmark (ranked third), and Universal-2 (October 2024, still on the API for cost-sensitive workloads) measures closer to 7–10% on real-world audio. AssemblyAI's own entity data (13.1% missed names) shows where "accurate" still breaks down — full data below.
WER (Word Error Rate) = (Substitutions + Deletions + Insertions) / total reference words — the NIST-standard ASR accuracy metric. Lower is better. Eight of the ten pages ranking for this question are written by AssemblyAI itself; every number below is labeled as a vendor claim or an independent measurement, with links in the Methodology & Sources section.
By VexaScribe Editorial · Published July 5, 2026 · Verified
VexaScribe does not sell a transcription API. This accuracy assessment has no commercial bias toward AssemblyAI. As of August 2026, we tested 14 speech-to-text models through their official APIs on 904 audio files across 16 standard datasets. WER via jiwer with lowercase + punctuation-stripped normalization; identical inputs per model.
AssemblyAI Accuracy in One Sentence
AssemblyAI is, by most independent evidence, one of the two most accurate commercial speech-to-text providers — peer-reviewed testing groups it with Whisper at the top on raw WER. Our own July 2026 benchmark of 903 files across 19 datasets ranks its Universal-3.5 Pro (July 2026) first among 9 models tested (7.0% aggregate WER). The honest caveats: its previous Universal-3 Pro model (deprecated July 2026) ranked third on Artificial Analysis's hardest benchmark subset, its clean-audio headline (1.52%) is roughly 4× better than its own real-world mean (5.6% across 26 datasets), and most tools built on AssemblyAI still call the older Universal-2 model rather than the current 3.5 Pro. All three caveats are covered below with sources.
Vendor Claims vs Independent Measurements
AssemblyAI publishes more of its own accuracy data than any competitor — including failure rates most vendors hide. That transparency deserves credit. It is still the company grading its own homework, so here is each headline claim next to what neutral sources measure.
| Metric | AssemblyAI's claim | Independent data | Context |
|---|---|---|---|
| Universal-3 Pro WER (deprecated July 2026) | 1.52% LibriSpeech clean; 5.6% mean across 26 real-world datasets | 2.3% on AgentTalk (AA-WER v2.0) — ranked 3rd. Novascribe did not test U-3 Pro directly (deprecated before our benchmark). | The vendor's own 1.52%-vs-5.6% spread is the honest headline: clean-audio numbers are ~4× better than its own real-world mean. U-3 Pro was superseded by U-3.5 Pro in July 2026 |
| Universal-2 WER (English) | “Industry-leading” across 99 languages | ~7–10% real-world | Consistent with Whisper Large-v3 (~8–12%) and Deepgram Nova-3 (~7–10%) — leading, but by 1–3 points, not a category apart |
| “Most accurate STT model” | AssemblyAI benchmarks page | Top-two in peer review; 3rd on latest AA index | arXiv 2408.16287 found AssemblyAI and Whisper the most accurate engines tested — the claim is close to true, but not uncontested |
| Missed Entity Rate (names) | 13.1% — “roughly half competitors’ rate” | No independent replication | Vendor-run but unusually honest: AssemblyAI publishes its own entity failure rates, which most vendors don't |
| Diarization speaker count | 2.9% error; phantom speakers −56% (streaming) | No independent replication | Vendor-run; directionally consistent with its strong reputation for built-in diarization |
| Universal-3.5 aggregate WER | Positioned as flagship model | 7.0% avg WER across 16 datasets (Novascribe July 2026) | Rank 3 of 14 tested (behind Speechmatics Melia-1 at 6.4% and Enhanced at 6.9%); #1 among promptable AI-transcription APIs — outperforms Deepgram Nova-3, Whisper-1, and GPT-4o on aggregate. Multilingual average 4.9% is category-leading; CommonVoice French 9.9% and Portuguese 12.1% are the weak spots |
Which AssemblyAI Model Are You Actually Using?
AssemblyAI shipped four model generations in 27 months — Universal-1 (April 2024), Universal-2 (October 2024), Universal-3 Pro (February 2026, deprecated July 2026), and Universal-3.5 Pro (July 2026 — the current flagship and the successor to U-3 Pro). Most third-party articles still describe Universal-3 Pro or Universal-2, and many production integrations call Universal-2 for cost reasons. If a tool "powered by AssemblyAI" underperforms the numbers on this page, check which generation it uses.
| Model | Released | Headline accuracy claim | Status |
|---|---|---|---|
| Universal-1 | April 2024 | 6.68% English WER (vendor) — the headline-WER generation | Superseded |
| Universal-2 | October 2024 | Built on Universal-1's WER; targeted proper nouns, formatting, alphanumerics — 73% blind human preference vs U-1 | Still on the API; recommended for cost-sensitive workloads |
| Universal-3 Pro | February 2026 | Promptable speech language model; 1.52% LibriSpeech clean, 5.6% mean across 26 real-world sets (vendor); 2.3% WER on Artificial Analysis AgentTalk (AA-WER v2.0), ranked third | Deprecated July 2026 — API rejects new requests |
| Universal-3.5 Pro | July 2026 | Current flagship successor to U-3 Pro; measured 7.0% aggregate WER across 16 datasets in Novascribe's July 2026 benchmark — rank 3 of 14 models (behind Speechmatics Melia-1 at 6.4% and Enhanced at 6.9%), and #1 among promptable AI-transcription APIs | Current flagship (what our benchmark tested) |
| Universal-3 Pro Streaming | 2026 | Real-time diarization, keyterm prompting, code-switching, 99+ languages | Voice-agent focused |
Sources: AssemblyAI's Universal-3 Pro announcement, Universal-2 release post, and Universal-3 Pro Streaming post. Verified July 5, 2026.
The U-3 / U-3.5 architectural shift matters more than the version numbers: this generation is a promptable speech language model — you can pass context ("this is a cardiology consult; expect drug names"), keyterms, and formatting instructions with the audio. Like Deepgram's keyterm prompting, this attacks the errors generic benchmarks don't measure: proper nouns, jargon, and domain terms. Universal-3.5 Pro inherits and extends this. Whisper offers no equivalent.
Where Universal-2 Lands on Standard Benchmarks
Cross-model WER on the eight standard English ASR test sets, compiled from the Hugging Face Open ASR Leaderboard and vendor documentation — the same numbers published on our Whisper and Deepgram accuracy pages. The Universal-3 Pro / Universal-3.5 Pro generation is too new to appear on the Open ASR Leaderboard; the closest external datapoint is 2.3% WER on AA-WER v2.0's AgentTalk subset (measured on U-3 Pro before its July 2026 deprecation). For our own July 2026 measurements on U-3.5 Pro, see the Novascribe benchmark section below.
| Benchmark | Domain | AssemblyAI Universal-2 | Whisper Large-v3 | Deepgram Nova-3 |
|---|---|---|---|---|
| LibriSpeech test-clean | Read English audiobook | 2.8% | 2.7% | 2.6% |
| LibriSpeech test-other | Read English, varied | 5.5% | 5.2% | 5.1% |
| TED-LIUM 3 | Conference talks | 3.9% | 4.0% | 3.6% |
| AMI (meeting headset) | Multi-speaker meetings | 14.1% | 15.9% | 13.4% |
| GigaSpeech | Diverse web English | 9.8% | 10.2% | 9.7% |
| Earnings-22 | Financial calls | 11.0% | 12.3% | 10.2% |
| CallHome | Conversational phone | 23.4% | 26.4% | 21.8% |
| CommonVoice 9 (English) | Crowdsourced diverse | 8.6% | 8.8% | 8.4% |
Novascribe's July 2026 Benchmark: Universal-3.5 vs 8 Alternatives
In July 2026 we ran 904 audio files across 16 standard benchmarks through 9 major speech-to-text models. AssemblyAI Universal-3.5 (which superseded the Universal-3 Pro numbers above) is the accuracy leader by our measurement — but with specific, honest weak spots you should know before deploying it.
AssemblyAI on English datasets
| Dataset | Domain | Universal-2 | Universal-3.5 | Rank |
|---|---|---|---|---|
| LibriSpeech test-clean | Audiobook read speech | 3.2% | 3.9% | 3rd–4th of 9 |
| AMI IHM | Multi-speaker meetings | 23.6% | 22.6% | 2nd of 9 |
| Earnings21 | Financial earnings calls | 13.5% | 12.4% | 2nd of 9 |
| TED-LIUM 3 | Long prepared speech | 4.2% | 4.8% | 3rd of 9 |
| GigaSpeech shard0 | Mixed web audio | 14.8% | 14.1% | 2nd of 9 |
| CommonVoice 9 EN | Crowdsourced diverse | 8.6% | n/a | 2nd (U-2) |
Ranks reflect position among the 9 AI-transcription APIs tested (AssemblyAI Universal-2 and Universal-3.5, Deepgram Nova-2, Nova-3 English, Nova-3 Multilingual and hosted Whisper Large, OpenAI Whisper-1, GPT-4o Transcribe, GPT-4o Mini Transcribe). Speechmatics' three tiers were added in expanded testing — Melia-1 leads aggregate WER at 6.4% across all 14 models tested; see the Speechmatics page. Universal-3.5 tied Nova-3 English or Nova-2 on several English datasets. WER via jiwer with lowercased, punctuation-stripped normalization.
AssemblyAI on multilingual (FLEURS clean vs CommonVoice accented)
| Language | FLEURS U-2 | FLEURS U-3.5 | CV U-2 | CV U-3.5 | Observation |
|---|---|---|---|---|---|
| German | 4.2% | 2.6% | 0.9% | 0.9% | Best of 9 clean + accented |
| French | 7.9% | 5.4% | 14.1% | 9.9% | Best clean; weak accented |
| Spanish | 1.5% | 1.9% | 4.6% | 2.9% | Best accented |
| Italian | 4.0% | 1.4% | 7.6% | 6.7% | Best single-language result |
| Portuguese | 3.4% | 4.7% | 16.1% | 12.1% | Weakest multilingual result |
Beyond WER: Where "Accurate" Breaks Down
A transcript can score 94% on WER and still misname every meeting attendee — names are a rounding error in word counts but the thing you actually search for. AssemblyAI is unusual in publishing its own entity-level failure rates, which makes an honest assessment possible. These are vendor-run numbers on Universal-3 Pro (measured before its July 2026 deprecation); AssemblyAI has not published equivalent per-metric figures for Universal-3.5 Pro. Treat them as best-case indicators of the U-3 generation family.
| Metric (Universal-3 Pro, vendor data) | Value | What it means |
|---|---|---|
| Missed Entity Rate — person/company names | 13.1% | Roughly 1 in 8 named entities still missed or misrendered — vendor-claimed to be about half competitors' rate |
| Missed Entity Rate — emails and URLs | 34.3% | 1 in 3 spoken emails/URLs wrong even on the flagship model — dictating addresses remains unreliable on every engine |
| Speaker count error (diarization) | 2.9% | Wrong number of detected speakers in ~3% of files |
| Phantom speaker reduction (streaming) | −56% | Universal-3 Pro Streaming vs prior streaming model |
| Medical entity error (Medical Mode) | 4.9% vs 7.3% | Universal-3 Pro Medical Mode vs competitors, vendor-run benchmark |
Source: assemblyai.com/benchmarks and the Universal-3 Pro Streaming announcement, accessed July 5, 2026. Novascribe's July 2026 benchmark did not measure entity-level error rates independently; the 13.1% missed-names figure remains vendor-reported. For aggregate WER and per-dataset cross-provider comparison, see the Novascribe 2026 Benchmark section above.
Accuracy by Audio Condition
What AssemblyAI's benchmark results translate to per audio scenario. Ranges centered on Universal-2 (what most integrations still call today). Universal-3.5 Pro extends this — in our own benchmark it beat Universal-2 on 11 of 15 datasets, with the biggest gains on multilingual and meetings.
| Audio Condition | Expected WER | Notes |
|---|---|---|
| Clean studio speech, 1 speaker | 3–5% | Podcasts, dictation, prepared speech |
| Conference talks | 3–4% | TED-LIUM-like audio |
| Conference call, 2 speakers | 7–10% | Business calls, decent microphones |
| Multi-speaker meetings (headset) | 13–16% | AMI benchmark: 14.1% (Universal-2) |
| Financial/jargon-heavy calls | 10–13% | Earnings-22: 11.0%; U-3 / U-3.5 promptable model reduces jargon misses vs U-2 |
| Conversational phone (8 kHz) | 20–26% | CallHome: 23.4% — hardest common scenario for every engine |
| Accented English | 8–14% | Top-two performer on non-native speech (arXiv 2408.16287) |
| Noisy / far-field audio | 15–25%+ | Degrades sharply; microphone quality dominates |
AssemblyAI vs Whisper vs Deepgram
The usual shortlist, on the axes that actually differ. Real-world WER from independent indexes; prices from vendor pricing pages, verified July 5, 2026.
| Engine | English WER | Entity handling | Price | Best for |
|---|---|---|---|---|
| AssemblyAI Universal-3.5 Pro (current) | 7.0% aggregate WER (Novascribe July 2026, rank 3 of 14; #1 promptable AI-transcription API) | Vendor has not published per-metric entity data yet | $0.21/hr base + $0.02/hr diarization | Max accuracy on real-world audio, multilingual, meetings, promptable + Audio Intelligence add-ons |
| Speechmatics Melia-1 (see /how-accurate-is-speechmatics) | 6.4% aggregate WER (Novascribe July 2026, best of 14 models) | No keyterm prompting or entity-error metrics published | $0.24/hr | Batch accuracy leader; not for real-time streaming |
| AssemblyAI Universal-3 Pro (deprecated July 2026) | 2.3% (AgentTalk, AA-WER v2.0) | 13.1% missed names (best published, U-3 Pro) | N/A — API rejects requests | Historical benchmark reference only |
| AssemblyAI Universal-2 | ~7–10% real-world; 11.6% aggregate (Novascribe) | Strong, pre-U3 baseline | $0.15/hr base + $0.02/hr diarization | 99-language batch, cost-sensitive workloads |
| Deepgram Nova-3 | ~7–10%; 12.3% English aggregate (Novascribe) | Keyterm prompting (100 terms) | $0.0043/min | Speed, telephony, cost per minute |
| Whisper Large-v3 | ~8–12% | No custom vocabulary support | Free (MIT, self-hosted) | Self-hosting, 99+ languages, budget |
| Whisper Large-v3-turbo | ~9–13% | No custom vocabulary support | Free (MIT, self-hosted) | Fast self-hosted pipelines |
Full Deepgram treatment — including why it wins on speed despite trailing on raw WER — on our Deepgram accuracy page.
When AssemblyAI Is the Right Choice — and When It Isn't
Choose AssemblyAI when:
- You need maximum accuracy on recorded audio among promptable AI-transcription APIs — Universal-3.5 Pro ranked rank 3 of 14 models in Novascribe's July 2026 benchmark (7.0% aggregate WER), #1 among APIs supporting keyterm/domain-prompt input. If pure batch WER is the only criterion and you don't need promptability, Speechmatics Melia-1 (6.4%) and Enhanced (6.9%) both edge it.
- Your audio is multilingual — U-3.5 Pro delivered 4.9% average WER across German, French, Spanish, Italian, Portuguese — best of every model tested
- Your audio is entity-heavy — names, companies, amounts — where the U-3 generation's published entity rates lead the industry
- You want built-in diarization that just works, including real-time speaker labels in streaming
- You can exploit prompting — passing domain context per request is the U-3 / U-3.5 generation's structural advantage
Look elsewhere when:
- You're cost-driven at volume — Deepgram undercuts it ($0.0043 vs $0.006/min) and Whisper is free to self-host
- You need the lowest streaming latency — Deepgram still owns the voice-agent latency benchmark
- You want full data control — there is no self-hosted AssemblyAI; Whisper runs air-gapped
- You don't write code — AssemblyAI is an API. There is no upload-a-file consumer product
Want top-tier accuracy without the API integration?
VexaScribe gives you Whisper Large-v3 accuracy through a simple upload interface — no code, from $2/mo. 100+ languages, speaker diarization, SRT/VTT/DOCX export.
Try VexaScribe FreeRelated Guides
Methodology & Sources
What WER actually measures
WER = (Substitutions + Deletions + Insertions) / Words in reference transcriptA WER of 5% means 95 of 100 reference words appear correctly. WER says nothing about which words are wrong — which is why this page also covers entity-level metrics (Missed Entity Rate) and diarization accuracy, where transcription quality is actually won or lost in practice.
Sources
- Universal-3 Pro announcement: assemblyai.com/blog/introducing-universal-3-pro (February 2026) — promptable speech language model architecture and pooled WER claims.
- Universal-3 Pro Streaming: announcement post — real-time diarization, phantom-speaker reduction (−56%), speaker-count error (2.9%).
- Universal-2 release: assemblyai.com/blog/universal-2 (October 2024) and Beyond Word Error Rate — 99-language coverage, Universal-1's 6.68% WER baseline, and the 73% blind human preference result.
- AssemblyAI benchmarks page: assemblyai.com/benchmarks — Missed Entity Rate data (13.1% names, 34.3% emails/URLs). Vendor-run.
- Artificial Analysis WER Index: artificialanalysis.ai/speech-to-text — 2.3% WER on AgentTalk (AA-WER v2.0), third-ranked; independent. AA-WER v2 weights: 50% AA-AgentTalk (conversational), 25% VoxPopuli (accented speech), 25% Earnings-22 (financial calls).
- Peer-reviewed evaluation: Measuring the Accuracy of Automatic Speech Recognition Solutions (arXiv 2408.16287) — AssemblyAI and Whisper ranked most accurate among tested engines.
- Hugging Face Open ASR Leaderboard: huggingface.co/spaces/hf-audio/open_asr_leaderboard — benchmark composite reference.
- AssemblyAI pricing: assemblyai.com/pricing — per-minute rates checked on the verification date.
Novascribe July 2026 Benchmark methodology
Test date: July 2026. 904 audio files across 16 standard benchmarks: LibriSpeech test-clean, AMI IHM, VoxConverse, Earnings21, TED-LIUM 3, GigaSpeech shard0, FLEURS (DE / FR / ES / IT / PT), CommonVoice 9 (DE / FR / ES / IT / PT), MLS-PT, plus 18 files of real Vexascribe production audio. 9 models tested through official APIs with identical inputs: AssemblyAI Universal-2 and Universal-3.5, Deepgram Nova-2, Nova-3 English, Nova-3 Multilingual and hosted Whisper Large, OpenAI Whisper-1, GPT-4o Transcribe, and GPT-4o Mini Transcribe. WER computed via jiwer with lowercase, punctuation-stripped normalization — the standard academic method. 95% bootstrap confidence intervals computed on datasets with ≥2 samples. No cherry-picking: all datasets included regardless of result; failures counted as errors.
Dataset licenses: LibriSpeech (CC BY 4.0), AMI (CC BY 4.0), VoxConverse (CC BY 4.0), Earnings21 (CC BY 4.0), TED-LIUM 3 (CC BY-NC-ND 3.0 — no transcripts reproduced), GigaSpeech (Apache 2.0), FLEURS (CC BY 4.0), CommonVoice (CC0), MLS (CC BY 4.0). Universal-3.5 is AssemblyAI's current shipping generation as tested; provider model versions update frequently, so results reflect performance at time of test.
Verification and update window
Published July 5, 2026. Refreshed with Novascribe July 2026 Benchmark data: July 15, 2026. Model versions tracked: AssemblyAI Universal-3 Pro / Universal-3.5 (February–July 2026), Universal-2 (October 2024), Universal-1 (April 2024), Deepgram Nova-3 (February 2025), Whisper Large-v3 (September 2023). Vendor claims, pricing, and benchmark numbers were cross-checked against the linked sources on the verification date. Where a claim has no independent replication, the page says so explicitly.
Frequently Asked Questions
What word error rate (WER) does AssemblyAI actually achieve?
Depends on the model and the audio. AssemblyAI's current flagship is Universal-3.5 Pro (July 2026, successor to the now-deprecated Universal-3 Pro). In Novascribe's July 2026 benchmark of 904 files across 16 datasets and 14 speech-to-text models, Universal-3.5 Pro achieved 7.0% aggregate WER — rank 3 of 14 overall (behind Speechmatics Melia-1 at 6.4% and Enhanced at 6.9%) and #1 among promptable AI-transcription APIs with prompt/keyterm support — with 4.9% average on multilingual audio (DE/FR/ES/IT/PT). Historically, Universal-3 Pro (February 2026, now deprecated) measured 2.3% WER on Artificial Analysis's AgentTalk (ranked third) with vendor-published numbers of 1.52% on LibriSpeech clean and 5.6% mean across 26 real-world sets. Universal-2 (October 2024, still on the API) measures 11.9% English AVG / 8.1% aggregate in our benchmark — about 3.2% on LibriSpeech, 23.6% on AMI meetings, 13.5% on Earnings21 financial calls. Clean-audio headlines run roughly 4× better than real-world means on every engine.
Is AssemblyAI more accurate than Whisper?
Yes, at the flagship tier. In Novascribe's July 2026 benchmark, Universal-3.5 Pro (7.0% aggregate WER) beat OpenAI Whisper-1 (8.3% aggregate) across 19 datasets. The gap was largest on multilingual (U-3.5: 4.9% avg vs Whisper-1: 6.7% avg). At the Universal-2 vs Whisper comparison, older peer-reviewed testing (arXiv 2408.16287) grouped AssemblyAI and Whisper together at the top and gaps were 1–3 percentage points. Whisper's counterweights: it's free to self-host under the MIT license, covers 99+ languages, and runs air-gapped. AssemblyAI's counterweights: built-in diarization, entity accuracy, promptable Universal-3.5 Pro.
Is AssemblyAI more accurate than Deepgram?
On aggregate recorded-audio accuracy, yes — Novascribe's July 2026 benchmark measured Universal-3.5 Pro at 7.0% aggregate WER vs Deepgram Nova-3 English at 8.9% and Nova-3 Multilingual at 8.2%. Multilingual is where AssemblyAI wins decisively: 4.9% average vs Nova-3 Multilingual's 8.2%. On specific English datasets they trade places — Nova-3 English wins AMI meetings by a hair (20.9% vs 22.6%) but Universal-3.5 Pro wins Earnings21 decisively (12.4% vs 18.1%). Practical rule: for maximum accuracy on batch transcription, AssemblyAI Universal-3.5 Pro leads; for streaming latency and price per minute ($0.0043 vs $0.006/min), Deepgram wins.
What is the difference between Universal-2, Universal-3 Pro, and Universal-3.5 Pro?
Universal-2 (October 2024) is a conventional ASR model covering 99 languages — still what most AssemblyAI integrations call today for cost reasons ($0.15/hr base). It prioritized proper nouns, formatting, and alphanumerics over headline WER. Universal-3 Pro (February 2026) was a promptable speech language model — you can pass domain context, keyterms, and formatting instructions alongside the audio. Universal-3 Pro is deprecated as of July 2026 — the API rejects new requests with 'universal-3-pro speech model(s) have been deprecated. Use speech_models: [universal-3-5-pro, universal-2] instead.' Universal-3.5 Pro (July 2026) is the current flagship successor — same promptable architecture, better multilingual accuracy. In our July 2026 benchmark, U-3.5 Pro won 11 of 15 head-to-head comparisons against U-2, with the biggest gains on multilingual (65% relative WER reduction on FLEURS Italian). If a tool 'powered by AssemblyAI' underperforms these numbers, check which model generation it actually uses.
How accurate is AssemblyAI's speaker diarization?
In Novascribe's July 2026 benchmark, Universal-2 and Universal-3.5 Pro produced identical diarization results — both correctly identified 4 speakers on 2 of 3 AMI meetings, got VoxConverse right on 16 of 30 files, and both undercounted by 2–5 speakers on large earnings calls with 5, 9, or 14 speakers. Diarization is effectively a tie between AssemblyAI's two current API models. AssemblyAI's own vendor numbers report a 2.9% speaker-count error rate and a 56% reduction in phantom speaker detections in the Universal-3 Pro Streaming variant. Note that speaker-count accuracy is not the same as word-level attribution accuracy: correctly counting two speakers doesn't guarantee every sentence is assigned to the right one.
How accurate is AssemblyAI on names, emails, and technical terms?
AssemblyAI publishes its own entity failure rates — rare transparency in this industry. The published Universal-3 Pro figures (measured before U-3 Pro's July 2026 deprecation) show 13.1% missed person/company names and 34.3% missed emails and URLs, which the company states is roughly half its competitors' error rate. AssemblyAI has not published equivalent per-metric figures for Universal-3.5 Pro; assume they're in the same range. Novascribe's benchmark did not measure entity-level errors independently — we tested word-level WER only. Read the vendor numbers both ways: best-in-class published entity accuracy, and still one wrong name in eight. If your use case depends on entities — legal, sales calls, journalism — test with your own audio and grade the names, not the overall word count.
Why do AssemblyAI's published numbers differ from independent benchmarks?
Benchmark shopping. AssemblyAI publishes benchmarks where AssemblyAI wins; Deepgram publishes benchmarks where Deepgram wins. Each vendor picks test sets, audio domains, and text normalization that flatter its model — the 1.52% headline comes from clean LibriSpeech audio (AssemblyAI's own 26-dataset real-world mean is 5.6%), while Artificial Analysis's uniform AA-WER v2 methodology (50% conversational AgentTalk, 25% accented VoxPopuli, 25% Earnings-22 financial calls) measured 2.3% with the model ranked third. None of these numbers is false. For fair comparisons, trust sources that run identical audio through every engine: Artificial Analysis, the Hugging Face Open ASR Leaderboard, and peer-reviewed studies like arXiv 2408.16287.
Does AssemblyAI handle accents and noisy audio well?
Among the best, but physics still applies. Peer-reviewed testing found AssemblyAI a top-two performer on non-native English speech. Expect roughly 8–14% WER on accented English, 13–16% on multi-speaker meetings, and 20–26% on conversational phone audio — degradation curves that apply to every engine, with AssemblyAI consistently near the top of the pack. Microphone quality and background noise remain bigger accuracy factors than engine choice once you're comparing the top three providers.
Which is more accurate, AssemblyAI Universal-3.5 or Deepgram Nova-3?
Universal-3.5 wins aggregate accuracy across our July 2026 benchmark: 7.0% average WER across 16 datasets vs Deepgram Nova-3 English's 8.9% and Nova-3 Multilingual's 8.2%. On specific English datasets they trade places — Nova-3 English wins AMI meetings by a hair (20.9% vs 22.6%) but Universal-3.5 wins Earnings21 decisively (12.4% vs 18.1%). On multilingual audio the gap is decisive: U-3.5 averaged 4.9% WER across DE/FR/ES/IT/PT while Nova-3 Multilingual averaged 8.2%. For English batch transcription they're close; for multilingual production audio Universal-3.5 wins clearly.
Does AssemblyAI handle accented multilingual audio well?
Mixed. In our July 2026 benchmark, Universal-3.5 was best-in-class on clean multilingual audio (FLEURS German 2.6%, Italian 1.4%, French 5.4%) and remarkable on accented German (CommonVoice-DE: 0.9% WER, actually better than clean audio — an inverse fragility pattern). But it struggled on accented French (CommonVoice-FR: 9.9%) and Portuguese (CommonVoice-PT: 12.1%) — its weakest multilingual results. For accented French, Deepgram's hosted Whisper Large actually outperforms U-3.5 (6.2% CommonVoice-FR) despite severe latency cost. If your production audio is heavily accented French or Portuguese, benchmark alternatives on your own audio before committing.
Is AssemblyAI Universal-3.5 Pro really 5-7% WER as the marketing suggests?
No, not on mixed real-world audio. AssemblyAI markets ~5–7% WER on selected benchmarks. In Novascribe's July 2026 benchmark of 6 English datasets and 904 audio files, Universal-3.5 Pro measured 11.6% average English WER — approximately 1.7–2.3× the vendor marketing range. This gap is structural across the industry: vendors report on their tuning distribution with favorable normalization; independent benchmarks average across a broader, harder audio mix with uniform normalization. Universal-3.5 Pro's aggregate 7.0% (which includes multilingual audio where it excels at 4.9%) is the number that beats most competitors. On English audio alone, expect ~11–12% WER — still strong, but not 5%.
How does AssemblyAI compare to Speechmatics Melia-1?
Speechmatics Melia-1 wins accuracy and price; AssemblyAI wins developer experience and Audio Intelligence add-ons. In Novascribe's July 2026 benchmark: Melia-1 aggregate WER 6.4% (rank 1 of 14) vs Universal-3.5 Pro 7.0% (rank 3). English AVG: Melia-1 10.5% vs Universal-3.5 Pro 11.6%. Multilingual: Melia-1 4.6% vs Universal-3.5 Pro 4.9%. Price: Melia-1 $0.24/hr vs Universal-3.5 Pro $0.30/hr. Where AssemblyAI wins: promptable architecture (keyterm boosting + domain context), rich Audio Intelligence bundle (summarization, sentiment, PII redaction, topic detection, chapters), cleaner SDKs (Python/Node/Go/Java), and streaming Universal-Streaming for real-time use cases (Speechmatics does not offer public streaming). Rule of thumb: Speechmatics for batch accuracy + cost; AssemblyAI for DX + streaming + Audio Intelligence.
Does AssemblyAI hallucinate like Whisper does?
Less frequently than Whisper. AssemblyAI's Universal series uses a purpose-built ASR architecture (not encoder-decoder like Whisper) which structurally reduces the specific 'invents plausible text during silence' hallucination pattern documented in Whisper (arXiv:2402.08021 measured ~1.4% of Whisper segments). Any neural ASR can substitute words on unclear or noisy audio — no engine is immune — but AssemblyAI's repetition-loop and silence-hallucination failure rates are materially lower than Whisper's. For safety-critical use (medical, legal), always keep human review regardless of engine.
Is Universal-3.5 Pro cheaper than Universal-2?
No, more expensive. Universal-2 costs approximately $0.15–$0.24/hr base depending on tier ($0.006–$0.10/min). Universal-3.5 Pro costs $0.30/hr base + $0.02/hr diarization ($0.005/min base). The premium buys promptable architecture (keyterms + domain context), better multilingual (4.9% avg vs Universal-2's 6.4% on our benchmark), and modest English gains (11.6% vs 11.9%). For pure English batch with no diarization needs where cost is the primary driver, Universal-2 remains a reasonable choice. For multilingual, meetings, or any workload that benefits from keyterm boosting, Universal-3.5 Pro pays for itself.