ASR & Speech Model Testing

Speech recognition testing across WER, CER, semantic and entity error rates, diarization and noisy-condition performance in 30+ languages and regional accents.

What We Test

Transcription accuracy; accent, dialect and demographic coverage; acoustic robustness; speaker attribution and structure; and downstream usability.

Why Word Error Rate Isn't Enough

Every error weighs the same, better capture can be penalised, the human reference is not neutral and valid formatting variants can score as errors.

MetricWhat it measuresWhen it matters
WERWord insertions, deletions and substitutionsBaseline comparability
CERCharacter-level errorIndic scripts and rich morphology
Semantic error rateWhether meaning survivedMeaning-sensitive products
Missed entity rateWhether critical terms survivedNames, drugs, amounts, IDs and phone numbers
Keyword / intent accuracyWhether downstream words survivedAssistants, IVR and commands
Diarization error rateSpeaker attribution and overlapCalls, meetings and consultations
Formatting accuracyCasing, punctuation and numeralsHuman or parser consumption
Latency / real-time factorSpeed under loadStreaming and agents

Accent, Dialect and Regional Testing

The external Voice of India benchmark evaluated 15 major Indian languages and 139 regional clusters—306,230 utterances, 536 hours and 36,691 speakers.

FindingPublished number
District-level WER range~4% Nainital to ~44% Mannarakkat
Best overall model, Hindi~5% WER
Same model, Bhojpuri20.9% WER
Same model, Maithili24.8% WER
Hindi-belt districtsBelow 10%
Kerala, interior KarnatakaSubstantially higher
Female vs male speakers3.1–4.3 points better for female speakers
Speakers aged 18–22Higher error rates
Audio-quality quartiles15.31% → 25.20% WER

Source: Voice of India benchmark, 2026. External evidence, not eQOURSE measurement.

Testing Under Real Conditions

ConditionWhat we testWhy it bites
Telephone bandwidth8 kHz, codecs and packet lossWideband-trained models degrade
Background noiseTraffic, crowd, home, office and machineryPerformance has a cliff
Reverberation & distanceEcho, far-field mic and speakerphoneCommon in real deployment
Overlapping speechInterruptions and multiple speakersDiarization and transcription fail together
Compression & pipelineCodec chains and resamplingProduction differs from test audio
Disfluency & hesitationFalse starts, fillers and correctionAbsent from read prompts

Voice Agents and Conversational Systems

Endpointing, barge-in, latency under load, error propagation, correction and recovery, plus ASR hallucination on silence, noise and non-speech.

The Reference Transcript Is Part of the Test

Native-variety transcribers, double-pass agreement, documented normalisation and a stated measurement ceiling.

Audio & Speech Annotation

How an Engagement Runs

StepActivityWindowOutput
01Scope & metric selectionWeek 1Metrics, languages and conditions
02Test set designWeeks 1–3Matched, stratified test set
03Reference transcriptionWeeks 2–4Native-variety references and agreement
04MeasurementWeeks 4–5Metrics per stratum
05Failure analysisWeeks 5–6Typed error backlog
06Report & walkthroughWeek 6Findings and reusable assets

What You Get

DeliverableContents
Accuracy reportEvery agreed metric by language, region, demographic and condition
Missed entity analysisCritical terms by category
Error taxonomyFailures ranked by frequency and impact
Audio examplesClip for every error class
Reference transcripts and test setVersioned and reusable
Agreement statisticsReference quality before findings
Condition curvesWhere performance falls off
Live walkthroughEngineering review

Where ASR Testing Goes Wrong

FailureWhat happensHow we avoid it
WER onlyCritical errors averaged awayUse entity and semantic metrics
Single national number4%–44% spread hiddenReport every stratum
Unmatched group difficultyCondition gap mistaken for accent gapControl content and condition
Studio audio onlyProduction path omittedReproduce the actual path
Standard-variety transcribersReference errors attributed to modelUse native-variety transcribers
No normalisation policyValid variants score as errorsDocument rules first
Read prompts onlyNatural speech absentInclude spontaneous conversation
Diarization ignoredRight words, wrong speakerMeasure DER
Hallucination untestedFabricated transcript passesUse silence and non-speech probes
Ceiling ignoredReference limits look like model errorState measurement ceiling

What We Do Not Do

We do not build or tune ASR models, report one blended number in isolation, accept unaudited sampling as representative or process voice data outside agreed controls.

How to Engage

ProgrammeBest fitWindow
Baseline accuracy programmeFull matched baseline5–6 weeks
Accent & dialect auditRegional coverage3–4 weeks
Release-cycle testingRepeat regression test1–2 weeks
Vendor comparisonMatched provider comparisonScoped

Related Services

AI Model Testing Human Evaluation & A/B Testing Audio & Speech Annotation AI Bias & Fairness Audit LLM Evaluation Data Collection Dataset QA & Label Audit

ASR & Speech Model Testing FAQs

What is ASR testing?

Structured measurement of a speech recognition system's accuracy and robustness — word and character error rates, plus semantic and entity-level error rates, diarization accuracy, formatting quality and latency — across the languages, accents, demographics and acoustic conditions your users actually present. The output is a per-stratum report and a typed error backlog, not a single accuracy figure.

Isn't WER enough?

No. WER weights every error equally, can penalise a more accurate model when references miss valid speech, and scores formatting variants as errors. We report WER for comparability alongside metrics that reflect what the product depends on.

What is missed entity rate?

The proportion of important terms — names, drug names, amounts, account IDs, phone numbers and addresses — that failed to survive transcription. Published comparisons reported 0% versus 8.3% on drug names and 19.6% versus 30.0% on phone numbers.

Why does regional accent matter so much in India?

The Voice of India benchmark covered 15 languages and 139 regional clusters across 306,230 utterances. District WER ranged roughly 4%–44%; one model scored about 5% on Hindi, 20.9% on Bhojpuri and 24.8% on Maithili.

Do you test code-mixed speech?

Yes. Hinglish, Tanglish, Benglish and mid-sentence switching are included as spontaneous speech, not only read prompts.

Can you test in noisy and telephone conditions?

Yes. We test telephone bandwidth, codec chains, graded noise, reverberation, far-field capture and overlapping speech, preferably from the real production path. Published Voice of India results moved from 15.31% to 25.20% WER across audio-quality quartiles.

How do you make sure reference transcripts are accurate?

We agree verbatim or clean standards, use native-variety transcribers, double-pass a sample, report inter-transcriber agreement, document normalisation and identify the human-reference ceiling.

What is ASR hallucination and do you test for it?

Some end-to-end speech models produce fluent fabricated text for silence, noise or non-speech. We probe it explicitly because fabricated speech can become fabricated intent downstream.

How is this different from audio annotation?

Audio and speech annotation produces transcripts and labelled audio used to train or fine-tune a model. ASR testing produces error rates and failure analysis used to decide whether a model is fit to ship.

How long does it take, and can we reuse the test set?

A first programme covering two or three languages with regional strata typically runs 5–6 weeks. The versioned test set is delivered to you, so later model versions can run in 1–2 weeks.

Find Out What Your Accuracy Number Is Hiding

Scope an ASR Test Programme