Speech Recognition and Text-to-Speech Data and Evaluation — Pashto, Dari, and 22 More Afghan Languages
Ariana Nexus builds and evaluates speech datasets for automatic speech recognition (ASR) and text-to-speech (TTS) systems in Pashto, Dari, and 22 more Afghan languages. We collect consented audio, transcribe and annotate it against a published standard, license voice talent for speech synthesis, and measure model accuracy — word error rate, character error rate, script fidelity and listener-rated quality — by dialect, speaker gender and recording condition.
Ariana Nexus is a Washington, D.C. consulting and professional services firm for the Afghan context, working across 24 Afghan languages including Pashto and Dari.
Why Afghan-language speech breaks the models you already have
The failure is not that accuracy is a little lower. For Pashto and Dari, published benchmarks show general-purpose speech models producing output that cannot be used at all — in the wrong script, under the wrong language label, with error rates above one hundred percent. These are third-party findings, not our marketing. They are the reason this service exists.
Zero-shot models do not transcribe Pashto — they guess
The output arrives in the wrong script
A score from one corpus does not predict the next one
The errors concentrate in specific sounds
Dari is labeled as Persian, and the label is wrong
When the data is right, the gap closes
Published Pashto word error rate, by condition
What we deliver
Four workstreams. They are sold separately and they compose. Most engagements begin with evaluation, because it is the cheapest way to find out what the data has to fix.
Speech data collection for ASR
- Read and scripted speech from phonetically balanced prompt sets
- Spontaneous and conversational speech, including two-party dialogue
- Telephony-band capture at 8 kHz and wideband at 16 kHz
- Far-field, in-vehicle, clinic and call-center acoustic conditions
- Controlled quotas by dialect, region, speaker gender and age band
- Domain prompt design: clinical, legal, financial, humanitarian, consumer
Voice data and licensing for TTS
- Professional Pashto and Dari voice talent, recorded to studio specification
- Phonetically and prosodically balanced scripts with question, statement and list contours
- Pronunciation adjudication for loanwords, personal names, place names and religious terms
- Licenses that name AI training and synthesis explicitly, with defined scope
- Withdrawal terms that propagate to derived models and to sublicensed buyers
- Synthesis provenance records for audio-marking obligations
Transcription, annotation and lexicon
- Verbatim orthographic transcription in Unicode Arabic script, normalized to a published spelling guide
- Utterance and word-level timestamps; forced alignment on request
- Speaker metadata: dialect, region, gender, age band, first and second languages
- Event tagging: noise, overlap, disfluency, truncation, unintelligible spans
- Code-switch tagging for Pashto–English and Dari–English
- Pronunciation lexicon and grapheme-to-phoneme rules with a documented phoneme inventory
Model evaluation and benchmarking
- Word error rate and character error rate on held-out sets, reported by slice
- Script fidelity rate and language-identification accuracy
- Phoneme-class error stratification against the documented inventory
- Named entity, numeral, date and currency accuracy
- ITU-T P.808 listening tests with native panels: ACR, DCR and CCR
- Intelligibility testing with semantically unpredictable sentences
Service summary
Where Pashto recognition errors concentrate
Delivery specification
The default specification. Every line is negotiable at scoping and fixed in the statement of work before recording begins.
How we measure a speech system
Evaluation is the part of this service that buyers underestimate and auditors ask about first. Every number we report carries the sample size, the slice it was measured on and a confidence interval, per ITU-T P.800.2 reporting practice.

Dialects are a specification, not a footnote
A corpus recorded entirely in one city produces a model that works in one city. Quotas are set at scoping and held; the delivered manifest reports what was actually captured against what was agreed.
The 24 Afghan languages
Pashto and Dari carry a standing bench. The remaining languages are scoped on notice, with a recruitment plan and a realistic timeline stated before the engagement is signed — not a fill guarantee we could not keep.
Iranian
Turkic
Indo-Aryan
Nuristani
Dravidian
Consent, licensing and provenance
A speech corpus is a file of human voices. In 2026 that makes it a regulated asset in three jurisdictions at once. The instruments below are the ones a buyer's counsel asks about; each produces an artifact we hand over with the data.
Where no benchmark exists, we build the one you can defend
There is no United States certification that tests any Afghan language, and no public benchmark that covers most of them. That absence is where this firm operates. We publish the rubric we grade against, we report agreement between graders, and we hand over the evidence rather than the claim.
Linguists, not a crowd
Afghanistan Dari is not Iranian Persian
Dialect and gender are quotas, not aspirations
Measurement is independent of production
One firm, one line of accountability
Jurisdiction and contributor safety
Documentation is a deliverable, not a favor
How an engagement runs
Scoping and specification
Baseline evaluation
Prompt and script design
Recruitment and consent
Capture and quality control
Annotation, lexicon and alignment
Delivery and documentation
Re-evaluation
The team behind this service
Ariana Nexus is operated by graduates and scholars of leading universities who work in these languages as their own. Speech data in Pashto and Dari cannot be staffed from a general annotation marketplace. It requires people who can hear the difference between a Kandahari and a Ningrahari retroflex, read an unfamiliar orthographic variant and judge whether it is an error or a regional spelling, and then defend that judgment in writing to a client's research team. That is the standard this team is held to.
Subject-matter depth is distributed deliberately: speech and machine-learning engineering, research methods and measurement, data governance and institutional law, and native linguistic authority in the varieties being recorded. No delivery is signed off by one discipline alone.
Hassan Ukasha

Zeba Haqbani

Hussain Ahmad

Wasil Peroz

Maryam Safi
How this team is deployed on an engagement
No delivery is signed off by one discipline alone. A corpus that clears the pipeline but not the linguistic review is not delivered, and a corpus that clears both but not the consent register is not delivered either.
Who buys this
Frontier and applied AI labs
Speech and voice technology companies
Platform and trust-and-safety teams
Product and localization teams
Federal programs and prime contractors
Universities and research groups
Four ways to engage
Questions buyers ask
What is ASR training data for Pashto and Dari?
It is recorded speech paired with a verified transcript and structured metadata, built so a speech recognition model can learn the sounds, words and spelling conventions of the language. For Pashto and Dari a usable corpus also needs dialect labels, a pronunciation lexicon and a written spelling standard, because both languages have real orthographic variation that a model will otherwise learn as noise.
How much audio do you need to fine-tune a speech recognition model for Pashto?
Published work shows meaningful movement from a few thousand utterances: one Pashto fine-tune on roughly 4,700 Common Voice clips reached 13.4% word error rate on that corpus's own test split. That is a starting point, not a product. Reaching accuracy that holds across dialects, telephony audio and spontaneous speech takes tens to hundreds of hours, and the mix matters more than the total.
Why do multilingual speech models fail on Pashto?
Three reasons compound. Pashto is largely absent from the pre-training corpora, so the model has no acoustic or lexical anchor. Its script shares a code block with Arabic, Persian and Urdu, so the decoder falls back to a language it does know. And its phoneme inventory contains retroflex and lateral fricative sounds that carry a disproportionate share of the error mass. The result is not slightly worse output; published benchmarks report zero-shot word error rates from 90% to 297%.
Is Dari the same as Persian or Farsi for speech data?
No. Afghanistan Dari and Iranian Persian differ in lexicon, phonology and register, and Afghan listeners hear the difference immediately. Iranian Persian voice talent recorded as Dari produces a system that sounds foreign to the users it was built for. We record Afghanistan Dari, label it as such, and say so in the datasheet.
Do you collect Hazaragi speech data?
Yes, with dialect-specific speakers. Hazaragi is a variety of Dari rather than a separate language, and that is how it is labeled in every manifest we deliver — counted among the 24 Afghan languages we cover and staffed separately, because a model trained only on Kabuli Dari will not serve Hazaragi speakers.
How do you evaluate a text-to-speech voice in Pashto or Dari?
With native listener panels run under ITU-T Recommendation P.808, the crowdsourced counterpart to laboratory P.800 testing. We run absolute category rating for naturalness, degradation and comparison category rating against a reference voice, and intelligibility testing with semantically unpredictable sentences. Scores are reported with the number of raters, the mean, the standard deviation and a confidence interval, following P.800.2 reporting practice. Automated proxies are used as screening, never as the finding.
What is word error rate, and what counts as good for a low-resource language?
Word error rate is the share of words a system inserts, deletes or substitutes against a reference transcript, so a rate above 100% is possible and common for Pashto. There is no universal threshold: what matters is the rate on your domain, your dialects and your acoustic conditions, compared against a baseline measured the same way. A single number from a public corpus is a marketing figure, not an engineering one.
How is speaker consent handled for voice data?
Consent is written, executed in the contributor's own language, and signed before recording — never retrofitted afterward. It names AI training and speech synthesis explicitly, states the retention term and the withdrawal channel, and is filed in a consent register delivered with the corpus. Withdrawal propagates to derived assets and to any sublicensed buyer.
Can we license a named voice for a Pashto or Dari TTS product?
Yes. Named-voice licenses state scope, term, territory, permitted product categories and withdrawal terms, and are drafted against state right-of-publicity law, the Tennessee ELVIS Act and the federal replica right proposed in the NO FAKES Act, so they do not need renegotiating if that bill passes.
Do you support telephony and far-field audio?
Yes. Telephony-band capture at 8 kHz, wideband at 16 kHz and studio at 48 kHz, recorded natively rather than downsampled to imitate a condition. Far-field, in-vehicle, clinic and call-center conditions are specified in the quota table at scoping.
What formats do you deliver, and will it drop into our pipeline?
WAV, 16-bit PCM, single channel, with transcripts and metadata in the manifest structure agreed at scoping. Directory layout, identifier scheme and field names are fixed before the first batch and held stable across the engagement, so integration is done once.
Which Afghan languages beyond Pashto and Dari can you collect?
All 24 in our coverage: the Iranian group including Balochi, Aimaq, Ormuri, Parachi and the Pamir languages; Turkic Uzbeki, Turkmeni and Kyrgyz; Indo-Aryan Pashayi, Gawarbati and Tirahi; the Nuristani group; and Brahui. For the smallest of these we state a recruitment plan and a realistic timeline at scoping rather than a fill guarantee, because several have only a few thousand speakers worldwide.
Can you evaluate a model we already trained, without building new data?
Yes, and it is where most engagements should start. Independent evaluation on held-out Afghan-language audio tells you which slices fail, whether the failure is acoustic, orthographic or lexical, and what data would actually move the number — before any collection budget is committed.
Do you work with United States federal agencies and prime contractors?
Yes. Ariana Nexus is registered for federal work and delivers with a United States footprint, documented provenance and records written for a contracting officer's review. We do not route data or inquiries through channels controlled by the de facto authorities in Afghanistan.
Terms on this page
Sources
Every third-party figure on this page is listed here with its source so it can be checked. Figures describing Ariana Nexus engagements are stated as commitments, not as measured results from another party's work.
- Radford, A. et al. Robust Speech Recognition via Large-Scale Weak Supervision (Whisper). 2023 — reports 99.0% word error rate for Whisper Base on Pashto in the FLEURS benchmark.
- Conneau, A. et al. FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech. 2022 — approximately 12 hours of speech per language across 102 languages.
- Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure and Cross-Domain Evaluation. 2026 — zero-shot word error rates of 90% to 297%; best zero-shot result 39.7%; character-class stratification of Pashto errors.
- PashtoTTS-Bench: automated screening for low-resource non-Latin-script text-to-speech. 2026 — language-identification failure on verified Pashto audio; automated intelligibility screening protocol.
- ML-SUPERB 2.0 — 141 languages across 15 corpora; Urdu at 21.8% character error rate on Common Voice against 56.9% on FLEURS.
- Pashto Common Voice corpus reports, 2026 — fine-tuned Pashto word error rate of 13.4% on the corpus test split.
- ITU-T Recommendation P.808 (06/2021). Subjective evaluation of speech quality with a crowdsourcing approach — ACR, DCR and CCR listening-test methods.
- ITU-T Recommendation P.800. Methods for subjective determination of transmission quality — the laboratory baseline for listening tests.
- ITU-T Recommendation P.800.2. Mean opinion score interpretation and reporting — the minimum information that must accompany a reported MOS.
- Regulation (EU) 2024/1689 (AI Act), Article 53(1)(d) and Article 50; European Commission AI Office, Explanatory Notice and Template for the Public Summary of Training Content, 24 July 2025.
- California Assembly Bill 2013, Generative Artificial Intelligence: Training Data Transparency — in force 1 January 2026.
- Illinois Biometric Information Privacy Act, 740 ILCS 14 — voiceprints as biometric identifiers.
- Tennessee Ensuring Likeness, Voice and Image Security (ELVIS) Act, 2024 — voice as a protected personal right.
Related services
LLM evaluation and benchmark development
Multilingual AI red teaming and safety testing
AI and machine translation quality review
All services
Start with a measurement
Send us a sample of your Afghan-language audio and the numbers you have today. We will tell you what the failure profile looks like, which slices are costing you accuracy and whether the answer is data, tuning or a different specification — before anyone proposes a corpus.