Skip to main content
HomeServicesSpeech recognition and text-to-speech data

Speech Recognition and Text-to-Speech Data and Evaluation — Pashto, Dari, and 22 More Afghan Languages

Ariana Nexus builds and evaluates speech datasets for automatic speech recognition (ASR) and text-to-speech (TTS) systems in Pashto, Dari, and 22 more Afghan languages. We collect consented audio, transcribe and annotate it against a published standard, license voice talent for speech synthesis, and measure model accuracy — word error rate, character error rate, script fidelity and listener-rated quality — by dialect, speaker gender and recording condition.

Ariana Nexus is a Washington, D.C. consulting and professional services firm for the Afghan context, working across 24 Afghan languages including Pashto and Dari.

99.0%
Word error rate reported for Whisper Base on Pashto, zero-shot, in the FLEURS benchmark.
Whisper (Radford et al., 2023)
90–297%
Range of zero-shot Pashto word error rates across ten multilingual speech models in a 2026 benchmark.
Benchmarking Multilingual Speech Models on Pashto (2026)
0.0%
Share of verified Pashto audio that one leading model labeled as Pashto in a 2026 screening study.
PashtoTTS-Bench / INSV-A (2026)
21.8% / 56.9%
The same model, the same language, two public corpora — character error rate on Common Voice against FLEURS.
ML-SUPERB 2.0
13.4%
Pashto word error rate after fine-tuning on roughly 4,700 Common Voice clips. The constraint is the corpus, not the architecture.
Pashto Common Voice corpus reports (2026)
These are third-party findings, not our marketing.
Every figure on this wall comes from published work by other people, listed in full at the foot of this page. We put them first because the case for this service is not an argument we made. It is a measurement someone else already took.

Why Afghan-language speech breaks the models you already have

The failure is not that accuracy is a little lower. For Pashto and Dari, published benchmarks show general-purpose speech models producing output that cannot be used at all — in the wrong script, under the wrong language label, with error rates above one hundred percent. These are third-party findings, not our marketing. They are the reason this service exists.

Zero-shot models do not transcribe Pashto — they guess

The original Whisper evaluation reports 99.0% word error rate for Whisper Base on Pashto in the FLEURS benchmark. A 2026 benchmark of ten multilingual speech models on Pashto reports zero-shot word error rates from 90% to 297%, with one configuration exceeding 400% through decoder looping. The best zero-shot result from any model in that study was 39.7%.
Source: Whisper (Radford et al., 2023); Benchmarking Multilingual Speech Models on Pashto (2026)

The output arrives in the wrong script

Multilingual models routinely render Pashto audio in Arabic, Persian or Urdu orthography. One 2026 screening study reports Whisper Large V3 returning a Pashto language label 0.0% of the time on verified Pashto audio. A transcript in the wrong script is not a degraded transcript. It is unusable output that passes automated quality gates.
Source: PashtoTTS-Bench / INSV-A screening (2026)

A score from one corpus does not predict the next one

ML-SUPERB 2.0, covering 141 languages across 15 corpora, reports per-language character error rate standard deviations of roughly 10% to 22% and places Urdu at 21.8% CER on Common Voice against 56.9% on FLEURS. Evaluation on a single public corpus tells you how a model performs on that corpus.
Source: ML-SUPERB 2.0

The errors concentrate in specific sounds

Character-class stratification of Pashto errors shows the retroflex series and the lateral fricatives carrying a disproportionate share of the error mass. More hours of undifferentiated audio does not fix a phoneme-coverage problem. Prompt design does.
Source: Benchmarking Multilingual Speech Models on Pashto (2026)

Dari is labeled as Persian, and the label is wrong

Afghanistan Dari and Iranian Persian differ in lexicon, phonology and register. Pipelines that merge them — or substitute Iranian voice talent for Afghan — produce systems that are fluent, confident and wrong in front of the users who most need them to be right.
Source: Ariana Nexus terminology standard

When the data is right, the gap closes

The same literature reports Pashto word error rate falling to 13.4% after fine-tuning on a Common Voice split, and to 35.1% with augmentation and no cross-domain degradation. The constraint is not the architecture. It is whether anyone built the corpus.
Source: Pashto Common Voice corpus reports (2026)

Published Pashto word error rate, by condition

Whisper Base, zero-shot (FLEURS)
99.0%
Ten-model benchmark, worst zero-shot
297%
Ten-model benchmark, best zero-shot
39.7%
Language-specific model (FLEURS)
34.6%
Augmented fine-tune, cross-domain
35.1%
Fine-tuned on in-corpus split
13.4%
Lower is better. Word error rate above 100% is possible because insertions are counted; it indicates output longer and less related to the reference than the reference itself. Figures are drawn from the published sources listed at the foot of this page, not from Ariana Nexus engagements.

What we deliver

Four workstreams. They are sold separately and they compose. Most engagements begin with evaluation, because it is the cheapest way to find out what the data has to fix.

Speech data collection for ASR

Consented recording in the varieties and conditions your product actually meets.
  • Read and scripted speech from phonetically balanced prompt sets
  • Spontaneous and conversational speech, including two-party dialogue
  • Telephony-band capture at 8 kHz and wideband at 16 kHz
  • Far-field, in-vehicle, clinic and call-center acoustic conditions
  • Controlled quotas by dialect, region, speaker gender and age band
  • Domain prompt design: clinical, legal, financial, humanitarian, consumer

Voice data and licensing for TTS

Studio recording and voice rights written so your legal team does not have to renegotiate them later.
  • Professional Pashto and Dari voice talent, recorded to studio specification
  • Phonetically and prosodically balanced scripts with question, statement and list contours
  • Pronunciation adjudication for loanwords, personal names, place names and religious terms
  • Licenses that name AI training and synthesis explicitly, with defined scope
  • Withdrawal terms that propagate to derived models and to sublicensed buyers
  • Synthesis provenance records for audio-marking obligations

Transcription, annotation and lexicon

The layer where most vendors quietly substitute a bilingual for a linguist.
  • Verbatim orthographic transcription in Unicode Arabic script, normalized to a published spelling guide
  • Utterance and word-level timestamps; forced alignment on request
  • Speaker metadata: dialect, region, gender, age band, first and second languages
  • Event tagging: noise, overlap, disfluency, truncation, unintelligible spans
  • Code-switch tagging for Pashto–English and Dari–English
  • Pronunciation lexicon and grapheme-to-phoneme rules with a documented phoneme inventory

Model evaluation and benchmarking

Independent measurement, with the test set held by someone who did not build the training data.
  • Word error rate and character error rate on held-out sets, reported by slice
  • Script fidelity rate and language-identification accuracy
  • Phoneme-class error stratification against the documented inventory
  • Named entity, numeral, date and currency accuracy
  • ITU-T P.808 listening tests with native panels: ACR, DCR and CCR
  • Intelligibility testing with semantically unpredictable sentences

Service summary

Languages
24 Afghan languages. Pashto and Dari at standing capacity; the remaining 22 on scoped notice.
Data types
Read, scripted, spontaneous and conversational speech; telephony at 8 kHz; far-field and in-vehicle capture.
Annotation
Verbatim orthographic transcription, speaker and dialect metadata, event and code-switch tags, forced alignment.
Voice for synthesis
Studio voice recording under licenses that name AI training explicitly, with scope limits and withdrawal propagation.
Evaluation
Word error rate, character error rate, script fidelity, language identification, and ITU-T P.808 listening tests with native panels.
Documentation
Every delivery ships with a datasheet, a consent register and a provenance file written for a procurement review.

Where Pashto recognition errors concentrate

Retroflex series ټ ډ ړ ڼ
30%
Lateral and post-velar fricatives ښ ږ
22%
Vowel length and zwarakay ə
16%
Named entities and numerals
14%
Code-switched English tokens
10%
Orthographic variants of the same word
8%
Indicative distribution used for prompt design and test-set construction. Each engagement produces its own measured distribution from the client's own audio.

Delivery specification

The default specification. Every line is negotiable at scoping and fixed in the statement of work before recording begins.

Parameter
Standard
Notes
Parameter
Audio format
Standard
WAV, PCM, 16-bit, single channel
Notes
FLAC on request; original session files retained
Parameter
Sample rates
Standard
48 kHz studio · 16 kHz wideband · 8 kHz telephony
Notes
Native capture, never upsampled to meet a specification
Parameter
Signal to noise
Standard
≥ 30 dB studio · ≥ 15 dB field
Notes
Measured and recorded per file, not per batch
Parameter
Transcription
Standard
Verbatim orthographic, Unicode Arabic script
Notes
Normalized against a written spelling guide delivered with the corpus
Parameter
Timestamps
Standard
Utterance level standard; word level on request
Notes
Forced alignment available for both
Parameter
Speaker metadata
Standard
Dialect, region, gender, age band, language history
Notes
Pseudonymous speaker identifiers; direct identifiers held separately
Parameter
Quality control
Standard
Independent second pass on a sampled share of every batch
Notes
Inter-annotator agreement reported with each delivery
Parameter
Lexicon
Standard
Pronunciation lexicon and grapheme-to-phoneme rules
Notes
Phoneme inventory documented per language and variety
Parameter
Packaging
Standard
Manifest, datasheet, consent register, provenance file
Notes
Directory structure agreed at scoping and held stable across batches

How we measure a speech system

Evaluation is the part of this service that buyers underestimate and auditors ask about first. Every number we report carries the sample size, the slice it was measured on and a confidence interval, per ITU-T P.800.2 reporting practice.

Measure
What it catches
Method
Measure
Word error rate (WER)
What it catches
Overall transcription accuracy
Method
Held-out test set, reported by dialect, gender and acoustic condition
Measure
Character error rate (CER)
What it catches
Morphological and affix errors a word metric hides
Method
Same test set, character-level alignment
Measure
Script fidelity rate
What it catches
Output rendered in Arabic, Persian or Urdu orthography instead of Pashto
Method
Character-class check against the documented inventory
Measure
Language identification
What it catches
Systems that never recognize the language is present
Method
Forced language-ID pass over verified audio
Measure
Phoneme-class stratification
What it catches
Which sounds carry the error mass
Method
Error alignment against the phoneme inventory
Measure
Entity and numeral accuracy
What it catches
Names, places, dates, currency and dosages
Method
Targeted slice with a curated entity list
Measure
Code-switch accuracy
What it catches
English tokens inside Pashto or Dari speech
Method
Tagged code-switch slice
Measure
Naturalness, MOS (ACR)
What it catches
Whether synthesized speech sounds like the language
Method
ITU-T P.808 absolute category rating, native panel, screened listeners
Measure
Preference, CMOS (DCR/CCR)
What it catches
Whether a new voice is better than the one shipping today
Method
ITU-T P.808 degradation and comparison category rating
Measure
Intelligibility
What it catches
Whether listeners can transcribe what the system said
Method
Semantically unpredictable sentences, listener transcription
Measure
Pronunciation adjudication
What it catches
Loanwords, names and religious terms said wrongly
Method
Native linguist review against a written rubric
An Afghan flag standing among United States flags on a university lawn
An Afghan flag among United States flags, Cornell University.

Dialects are a specification, not a footnote

A corpus recorded entirely in one city produces a model that works in one city. Quotas are set at scoping and held; the delivered manifest reports what was actually captured against what was agreed.

Language
Varieties we staff
Why it matters for speech
Language
Pashto
Varieties we staff
Kandahari (southern), Ningrahari and eastern, Wardak and central, Waziri
Why it matters for speech
The retroflex and fricative series differ audibly across these varieties; a model trained on one mislabels the others
Language
Dari
Varieties we staff
Kabuli, Herati, Badakhshani, Mazari, and Hazaragi as a variety of Dari
Why it matters for speech
Vowel realization and lexical choice separate Afghanistan Dari from Iranian Persian, which most pipelines substitute silently
Language
Uzbeki
Varieties we staff
Afghan Uzbeki of the northern provinces
Why it matters for speech
Distinct from the standard Uzbek of Uzbekistan in phonology and loan vocabulary
Language
Turkmeni
Varieties we staff
Afghan Turkmeni
Why it matters for speech
Distinct from the standard Turkmen of Turkmenistan; substitution is common and detectable
Language
Pashayi, Nuristani group
Varieties we staff
Regional varieties by valley
Why it matters for speech
Small speaker populations with no standardized orthography; capture requires a written convention agreed first

The 24 Afghan languages

Pashto and Dari carry a standing bench. The remaining languages are scoped on notice, with a recruitment plan and a realistic timeline stated before the engagement is signed — not a fill guarantee we could not keep.

Iranian

13
Pashto
پښتو
Dari
دری
Hazaragi
هزارگی
Aimaq
ایماق
Balochi
بلوچی
Ormuri
اورموړی
Parachi
پراچی
Wakhi
وخی
Shughni
شغنی
Sanglechi
سنگلیچی
Ishkashimi
اشکاشمی
Munji
منجی
Yidgha
یدغه

Turkic

3
Uzbeki
اوزبیکی
Turkmeni
ترکمنی
Kyrgyz
قرغیزی

Indo-Aryan

3
Pashayi
پشه‌یی
Gawarbati
گواربتی
Tirahi
تیراهی

Nuristani

4
Nuristani (Ashkun group)
نورستانی
Kati
کتی
Prasun
پارون
Waigali
وایگلی

Dravidian

1
Brahui
براهوی
Hazaragi is counted among the 24 and staffed with dialect-specific speakers; in linguistic terms it is a variety of Dari, and it is labeled that way in every manifest we deliver.

Consent, licensing and provenance

A speech corpus is a file of human voices. In 2026 that makes it a regulated asset in three jurisdictions at once. The instruments below are the ones a buyer's counsel asks about; each produces an artifact we hand over with the data.

Instrument
What it requires
What we produce
Instrument
EU AI Act, Article 53(1)(d)
What it requires
A public summary of training content, on the AI Office template. The duty has applied since 2 August 2025; the Commission's enforcement powers became active on 2 August 2026, and models placed on the market before 2 August 2025 must publish by 2 August 2027.
What we produce
Dataset-level source, characteristics, volume and processing records mapped to the template's fields
Instrument
EU AI Act, Article 50
What it requires
Machine-readable marking of synthetic audio, applying from 2 August 2026.
What we produce
A synthesis provenance record for every licensed voice and generated sample
Instrument
California AB 2013
What it requires
A public, high-level training-data summary for generative systems made available in California, in force since 1 January 2026 — sources and ownership, volume, collection method, personal-information status and synthetic-data use.
What we produce
A per-corpus disclosure pack written to the statute's enumerated fields
Instrument
GDPR and UK GDPR, Article 9
What it requires
Voice recordings used for identification are biometric data; explicit consent, withdrawal and erasure apply.
What we produce
Signed consent in the contributor's own language, a withdrawal channel, and deletion that propagates
Instrument
Illinois Biometric Information Privacy Act
What it requires
Written consent before a voiceprint is captured, with a published retention schedule.
What we produce
Consent executed before recording, never retrofitted, with the retention term on the face of the document
Instrument
Tennessee ELVIS Act and state right-of-publicity law
What it requires
Voice is a protected personal right; named-voice replication requires a license with defined scope.
What we produce
Named-voice licenses with scope, term, territory and withdrawal terms stated
Instrument
NO FAKES Act (introduced, not enacted)
What it requires
A proposed federal right against unauthorized AI replicas of voice and likeness.
What we produce
Contributor agreements drafted to survive passage without renegotiation
Ariana Nexus is not a law firm and this page is not legal advice. We produce the records your counsel needs and we work alongside them, never in place of them.

Where no benchmark exists, we build the one you can defend

There is no United States certification that tests any Afghan language, and no public benchmark that covers most of them. That absence is where this firm operates. We publish the rubric we grade against, we report agreement between graders, and we hand over the evidence rather than the claim.

Linguists, not a crowd

Every contributor is recruited, trained, paid and recorded by name in the consent register. No anonymous marketplace labor sits anywhere in the pipeline, because a corpus you cannot trace is a corpus you cannot disclose.

Afghanistan Dari is not Iranian Persian

We refuse the substitution that most pipelines make silently, in voice talent, in transcription and in labeling. The same discipline applies to Afghan Uzbeki against standard Uzbek and Afghan Turkmeni against standard Turkmen.

Dialect and gender are quotas, not aspirations

The delivered manifest reports what was captured against what was agreed, variety by variety. Where we miss a quota, the manifest says so.

Measurement is independent of production

The team that holds the test set is not the team that built the training data. That separation is the reason our numbers survive a client's own replication.

One firm, one line of accountability

Collection, transcription, lexicon, voice licensing and evaluation are produced in-house. No brokered specialists, no subcontracted annotation floor, no vendor chain to audit.

Jurisdiction and contributor safety

Delivery is United States-based. Ariana Nexus does not route documents, data or inquiries through channels controlled by the de facto authorities in Afghanistan, and contributor identities are held separately from the delivered corpus.

Documentation is a deliverable, not a favor

Datasheet, consent register, provenance file and evaluation report ship with the data. They are written for the reviewer who arrives two years later, not for the buyer who is already convinced.

How an engagement runs

01

Scoping and specification

We agree the varieties, the acoustic conditions, the quota table and the acceptance thresholds in writing. Nothing is recorded before the specification is signed.
1–2 weeks
02

Baseline evaluation

We measure what your current system does on Afghan-language audio, by slice, and publish the failure profile. This is the number every later delivery is measured against.
2–3 weeks
03

Prompt and script design

Prompt sets are built for phoneme coverage, entity density and the domains your product serves, then reviewed by native linguists before a speaker is booked.
2 weeks
04

Recruitment and consent

Contributors are recruited against the quota table, paid, and consented in their own language before capture. The consent register opens here and stays open.
Continuous
05

Capture and quality control

Recording runs in batches. Every batch gets an independent second pass on a sampled share, with inter-annotator agreement reported before the batch is accepted.
By volume
06

Annotation, lexicon and alignment

Transcription, metadata, event tagging and lexicon work run against the written standard, with linguist adjudication on disputed items.
By volume
07

Delivery and documentation

Audio, transcripts, metadata, lexicon, datasheet, consent register and provenance file arrive together. Incomplete documentation is a failed delivery.
Per batch
08

Re-evaluation

The same test protocol is re-run after training. You receive the movement, by slice, with the sample size and the confidence interval.
Post-training

The team behind this service

Ariana Nexus is operated by graduates and scholars of leading universities who work in these languages as their own. Speech data in Pashto and Dari cannot be staffed from a general annotation marketplace. It requires people who can hear the difference between a Kandahari and a Ningrahari retroflex, read an unfamiliar orthographic variant and judge whether it is an error or a regional spelling, and then defend that judgment in writing to a client's research team. That is the standard this team is held to.

Subject-matter depth is distributed deliberately: speech and machine-learning engineering, research methods and measurement, data governance and institutional law, and native linguistic authority in the varieties being recorded. No delivery is signed off by one discipline alone.

Hassan Ukasha

Managing Partner
Executive accountability for the firm's operations and for this program
B.S., Cornell University · M.P.H., Cornell University
Hassan Ukasha carries executive accountability for this service line: its scope, its staffing, the standards it is held to, and the evidence attached to every claim on this page. Engagements of institutional significance are reviewed by him before a specification is signed.
Zeba Haqbani

Zeba Haqbani

Senior Partner
Speech data platform, pipeline engineering and delivery systems
B.A., The American University of Afghanistan · B.A., University of British Columbia
Owns the data platform: capture tooling, manifest integrity, alignment pipelines and the packaging that makes a corpus drop into a training run without a week of cleanup.
Hussain Ahmad

Hussain Ahmad

Principal
Model evaluation, benchmarking and error analysis
M.Eng., Cornell University · Ph.D., University of Chicago
Holds the test sets. Designs the slices, sets the sampling frame and the acceptance thresholds, runs the word error rate and stratification passes, and writes the evaluation report the client's own researchers will try to replicate.
Wasil Peroz

Wasil Peroz

Principal
Data governance, consent instruments and licensing
B.A., Milli University · M.Sc., Otto-von-Guericke University Magdeburg
Drafts the contributor consent and voice licenses, maintains the provenance file, and maps each corpus to the disclosure obligations the client will be measured against.
Maryam Safi

Maryam Safi

Principal
Linguistic authority, orthography and pronunciation adjudication
B.A., Cornell University
Rules on disputed transcriptions, maintains the spelling guide and the pronunciation lexicon, and adjudicates loanwords, personal names and religious terminology.

How this team is deployed on an engagement

Stage
Who owns it
What they sign off
Stage
Specification and sampling
Who owns it
Hussain Ahmad
What they sign off
The quota table, the acceptance thresholds and the held-out test set
Stage
Capture and pipeline
Who owns it
Zeba Haqbani
What they sign off
Manifest integrity, alignment output and delivery packaging
Stage
Consent, licensing and provenance
Who owns it
Wasil Peroz
What they sign off
The consent register, the voice licenses and the disclosure pack
Stage
Transcription and linguistic adjudication
Who owns it
Maryam Safi
What they sign off
The spelling guide, the pronunciation lexicon and every disputed item

No delivery is signed off by one discipline alone. A corpus that clears the pipeline but not the linguistic review is not delivered, and a corpus that clears both but not the consent register is not delivered either.

Who buys this

Frontier and applied AI labs

You need Pashto and Dari coverage you can disclose under the training-content summary, and a benchmark that is not a public corpus your competitors also trained on.

Speech and voice technology companies

You are shipping ASR or TTS into a market with Afghan users and your current numbers came from a corpus of a few hours of read speech.

Platform and trust-and-safety teams

You moderate Afghan-language audio and video and need transcription accuracy you can measure, in the varieties your users actually speak.

Product and localization teams

You are adding a voice interface, an IVR path or captioning and need pronunciation, prosody and text-normalization judgments made by someone accountable for them.

Federal programs and prime contractors

You are buying Pashto or Dari speech capability and need a supplier with documented provenance, a United States delivery footprint and records that survive a contracting officer's review.

Universities and research groups

You are publishing on low-resource speech and need a corpus with a datasheet, an ethics trail and reproducible evaluation.

Four ways to engage

Model
What it is
Typical output
Choose it when
Model
Evaluation only
What it is
Independent measurement of a system you already have
Typical output
Failure profile by slice, script-fidelity and language-ID findings, prioritized data plan
Choose it when
You need to know what is wrong before you spend on data
Model
Pilot corpus
What it is
A bounded first corpus in one or two varieties
Typical output
Audio, transcripts, metadata, lexicon seed, datasheet, before-and-after evaluation
Choose it when
You need evidence that a corpus moves your numbers before committing a budget
Model
Full corpus build
What it is
Production collection against a signed specification
Typical output
Batched delivery with agreement reporting, full documentation pack, re-evaluation
Choose it when
The pilot cleared and the model needs volume and variety coverage
Model
Standing collection
What it is
A continuing bench with monthly capacity
Typical output
Recurring batches, drift monitoring, refreshed test sets, quarterly evaluation
Choose it when
Your product ships continuously and the data has to keep up with it

Questions buyers ask

What is ASR training data for Pashto and Dari?

It is recorded speech paired with a verified transcript and structured metadata, built so a speech recognition model can learn the sounds, words and spelling conventions of the language. For Pashto and Dari a usable corpus also needs dialect labels, a pronunciation lexicon and a written spelling standard, because both languages have real orthographic variation that a model will otherwise learn as noise.

How much audio do you need to fine-tune a speech recognition model for Pashto?

Published work shows meaningful movement from a few thousand utterances: one Pashto fine-tune on roughly 4,700 Common Voice clips reached 13.4% word error rate on that corpus's own test split. That is a starting point, not a product. Reaching accuracy that holds across dialects, telephony audio and spontaneous speech takes tens to hundreds of hours, and the mix matters more than the total.

Why do multilingual speech models fail on Pashto?

Three reasons compound. Pashto is largely absent from the pre-training corpora, so the model has no acoustic or lexical anchor. Its script shares a code block with Arabic, Persian and Urdu, so the decoder falls back to a language it does know. And its phoneme inventory contains retroflex and lateral fricative sounds that carry a disproportionate share of the error mass. The result is not slightly worse output; published benchmarks report zero-shot word error rates from 90% to 297%.

Is Dari the same as Persian or Farsi for speech data?

No. Afghanistan Dari and Iranian Persian differ in lexicon, phonology and register, and Afghan listeners hear the difference immediately. Iranian Persian voice talent recorded as Dari produces a system that sounds foreign to the users it was built for. We record Afghanistan Dari, label it as such, and say so in the datasheet.

Do you collect Hazaragi speech data?

Yes, with dialect-specific speakers. Hazaragi is a variety of Dari rather than a separate language, and that is how it is labeled in every manifest we deliver — counted among the 24 Afghan languages we cover and staffed separately, because a model trained only on Kabuli Dari will not serve Hazaragi speakers.

How do you evaluate a text-to-speech voice in Pashto or Dari?

With native listener panels run under ITU-T Recommendation P.808, the crowdsourced counterpart to laboratory P.800 testing. We run absolute category rating for naturalness, degradation and comparison category rating against a reference voice, and intelligibility testing with semantically unpredictable sentences. Scores are reported with the number of raters, the mean, the standard deviation and a confidence interval, following P.800.2 reporting practice. Automated proxies are used as screening, never as the finding.

What is word error rate, and what counts as good for a low-resource language?

Word error rate is the share of words a system inserts, deletes or substitutes against a reference transcript, so a rate above 100% is possible and common for Pashto. There is no universal threshold: what matters is the rate on your domain, your dialects and your acoustic conditions, compared against a baseline measured the same way. A single number from a public corpus is a marketing figure, not an engineering one.

How is speaker consent handled for voice data?

Consent is written, executed in the contributor's own language, and signed before recording — never retrofitted afterward. It names AI training and speech synthesis explicitly, states the retention term and the withdrawal channel, and is filed in a consent register delivered with the corpus. Withdrawal propagates to derived assets and to any sublicensed buyer.

Can we license a named voice for a Pashto or Dari TTS product?

Yes. Named-voice licenses state scope, term, territory, permitted product categories and withdrawal terms, and are drafted against state right-of-publicity law, the Tennessee ELVIS Act and the federal replica right proposed in the NO FAKES Act, so they do not need renegotiating if that bill passes.

Do you support telephony and far-field audio?

Yes. Telephony-band capture at 8 kHz, wideband at 16 kHz and studio at 48 kHz, recorded natively rather than downsampled to imitate a condition. Far-field, in-vehicle, clinic and call-center conditions are specified in the quota table at scoping.

What formats do you deliver, and will it drop into our pipeline?

WAV, 16-bit PCM, single channel, with transcripts and metadata in the manifest structure agreed at scoping. Directory layout, identifier scheme and field names are fixed before the first batch and held stable across the engagement, so integration is done once.

Which Afghan languages beyond Pashto and Dari can you collect?

All 24 in our coverage: the Iranian group including Balochi, Aimaq, Ormuri, Parachi and the Pamir languages; Turkic Uzbeki, Turkmeni and Kyrgyz; Indo-Aryan Pashayi, Gawarbati and Tirahi; the Nuristani group; and Brahui. For the smallest of these we state a recruitment plan and a realistic timeline at scoping rather than a fill guarantee, because several have only a few thousand speakers worldwide.

Can you evaluate a model we already trained, without building new data?

Yes, and it is where most engagements should start. Independent evaluation on held-out Afghan-language audio tells you which slices fail, whether the failure is acoustic, orthographic or lexical, and what data would actually move the number — before any collection budget is committed.

Do you work with United States federal agencies and prime contractors?

Yes. Ariana Nexus is registered for federal work and delivers with a United States footprint, documented provenance and records written for a contracting officer's review. We do not route data or inquiries through channels controlled by the de facto authorities in Afghanistan.

Terms on this page

Term
ASR
Definition
Automatic speech recognition — converting recorded speech into text.
Term
TTS
Definition
Text-to-speech — synthesizing spoken audio from written text.
Term
WER
Definition
Word error rate — insertions, deletions and substitutions against a reference transcript, as a share of reference words.
Term
CER
Definition
Character error rate — the same measure at character level, which exposes affix and morphology errors a word metric hides.
Term
MOS
Definition
Mean opinion score — listener-rated quality on a five-point scale, collected under ITU-T P.800 or P.808.
Term
Script fidelity
Definition
Whether output is rendered in the language's own orthography rather than a neighboring script.
Term
Code-switching
Definition
Alternating between languages inside a single utterance, common in Pashto–English and Dari–English speech.
Term
Forced alignment
Definition
Automatically matching a transcript to its audio in time, producing word-level boundaries.
Term
G2P
Definition
Grapheme-to-phoneme — rules converting written forms into pronunciations, required for both recognition and synthesis.
Term
Datasheet
Definition
The document describing a dataset's composition, collection method, consent basis and intended use.

Sources

Every third-party figure on this page is listed here with its source so it can be checked. Figures describing Ariana Nexus engagements are stated as commitments, not as measured results from another party's work.

  1. Radford, A. et al. Robust Speech Recognition via Large-Scale Weak Supervision (Whisper). 2023 — reports 99.0% word error rate for Whisper Base on Pashto in the FLEURS benchmark.
  2. Conneau, A. et al. FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech. 2022 — approximately 12 hours of speech per language across 102 languages.
  3. Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure and Cross-Domain Evaluation. 2026 — zero-shot word error rates of 90% to 297%; best zero-shot result 39.7%; character-class stratification of Pashto errors.
  4. PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech. 2026 — language-identification failure on verified Pashto audio; automated intelligibility screening protocol.
  5. ML-SUPERB 2.0 — 141 languages across 15 corpora; Urdu at 21.8% character error rate on Common Voice against 56.9% on FLEURS.
  6. Pashto Common Voice corpus reports, 2026 — fine-tuned Pashto word error rate of 13.4% on the corpus test split.
  7. ITU-T Recommendation P.808 (06/2021). Subjective evaluation of speech quality with a crowdsourcing approach — ACR, DCR and CCR listening-test methods.
  8. ITU-T Recommendation P.800. Methods for subjective determination of transmission quality — the laboratory baseline for listening tests.
  9. ITU-T Recommendation P.800.2. Mean opinion score interpretation and reporting — the minimum information that must accompany a reported MOS.
  10. Regulation (EU) 2024/1689 (AI Act), Article 53(1)(d) and Article 50; European Commission AI Office, Explanatory Notice and Template for the Public Summary of Training Content, 24 July 2025.
  11. California Assembly Bill 2013, Generative Artificial Intelligence: Training Data Transparency — in force 1 January 2026.
  12. Illinois Biometric Information Privacy Act, 740 ILCS 14 — voiceprints as biometric identifiers.
  13. Tennessee Ensuring Likeness, Voice and Image Security (ELVIS) Act, 2024 — voice as a protected personal right.

Start with a measurement

Send us a sample of your Afghan-language audio and the numbers you have today. We will tell you what the failure profile looks like, which slices are costing you accuracy and whether the answer is data, tuning or a different specification — before anyone proposes a corpus.

Engagements begin under a mutual non-disclosure agreement. Washington, D.C. · United States and worldwide.
Ariana Nexus · 1717 Pennsylvania Avenue NW, 10th Floor · Washington, D.C. 20006