Skip to main content

Pashto and Dari Training Data Collection and Annotation: Text, Speech and Preference Data

Ariana Nexus collects and annotates Pashto and Dari training data for AI models: instruction and chat data, spoken conversation, transcription, and human preference data for RLHF and DPO. Native Afghan linguists write the guidelines, recruit and consent every contributor, and document every dataset for audit.

Languages
Pashto and Dari, with 22 more Afghan languages scoped per program
Data
Text, speech and preference data, and annotation of data you already hold
Delivered as
JSONL, Parquet, WAV or FLAC and TextGrid, with a datasheet for every release
Engagements begin with
A paid pilot, scored against written acceptance thresholds
Headquarters
Washington, D.C.

What Pashto and Dari training data collection and annotation includes

Pashto and Dari training data collection and annotation is the work of creating, recording and labeling the examples an AI model learns from, in the two main languages of Afghanistan. It covers written text (instructions, conversations, domain corpora and parallel translations), speech (conversations, spoken prompts and transcripts) and preference data: human judgments of which model response is better, and why.

Who commissions Pashto and Dari training data

  • Frontier AI labs and model developers adding Pashto and Dari
  • Speech and voice AI teams building assistants and speech-to-speech models
  • Trust and safety teams training classifiers for Pashto and Dari content
  • Machine translation teams that need Afghan-language parallel data
  • Federal agencies and prime contractors building Pashto and Dari language technology
  • AI data companies that need a native Pashto and Dari bench
  • Universities and research institutions working on low-resource languages

Pashto and Dari training data we collect and annotate

Three families of data, each built to a written specification and annotated by native linguists. Every annotation service can also be applied to data you already hold.

Text

Pashto and Dari text data

Written data for supervised fine-tuning, chat, translation and classification, written natively in Pashto and Dari rather than machine-translated from English.
Instruction and chat data (SFT)
Prompts and ideal responses written by native speakers across the tasks your model must perform.
Multi-turn conversations
Realistic dialogues in everyday, health, legal, public-service, education and commercial settings.
Domain text corpora
Licensed, consented writing in the registers your model must handle: formal, colloquial, broadcast and online.
Parallel corpora
Pashto–English, Dari–English and Pashto–Dari sentence pairs for machine translation, aligned and reviewed.
Text annotation
Named entities, intent and slots, sentiment, topic, toxicity and safety labels, under written guidelines.
OCR ground truth
Printed and handwritten Pashto and Dari in Naskh and Nastaliq styles, transcribed line by line.
Romanized and code-switched text
Pashto and Dari as people actually type them, in Latin script and mixed with English, normalized and labeled.
Speech

Pashto and Dari speech data

Spoken data for voice assistants, speech-to-speech models and spoken-language understanding, recorded with informed consent.
Conversational speech
Two-speaker dialogues and spontaneous monologues, with speaker panels balanced by dialect, region, gender and age.
Spoken prompts
Questions and commands for voice assistants and speech-to-speech models, spoken the way users actually ask.
Phone, wideband and far-field capture
8 kHz telephone, 16 kHz wideband and room-distance recordings for the channel your product runs on.
Transcription
Verbatim or clean transcripts with timestamps, speaker turns and non-speech events, in normalized Pashto or Dari script.
Speech annotation
Speaker diarization, dialect and accent labels, emotion and intent tags on audio.
Speech translation pairs
Pashto or Dari audio aligned to English translations for speech translation models.
Preference

Pashto and Dari preference data for RLHF and DPO

Human judgments that teach a model which response is better in Pashto and Dari: the data behind reward models, RLHF and DPO.
Pairwise preferences
Native raters choose between two model responses and write the reason, in English or the source language.
Rankings and rubric ratings
Responses scored on helpfulness, accuracy, safety, dialect fit and cultural fit.
Rewrites
Native speakers correct or rewrite a response to produce the preferred answer for DPO and SFT.
Safety preference data
Refusals and safe completions judged by raters who understand the local context of harm.
Multi-turn conversation ratings
Whole conversations rated turn by turn, including tone and register across the exchange.
Spoken-response preference
Voice model replies judged for pronunciation, dialect and naturalness by native listeners.
Rubrics for reinforcement learning
Expert-written grading rubrics that score Pashto and Dari responses for rubric-based rewards.

What Pashto and Dari AI models are still missing

Read speech is no longer the bottleneck for Pashto: Mozilla Common Voice now holds more than 3,000 validated hours of it. What public data still lacks is natural conversation, domain language, dialect and gender balance, Afghan Dari as distinct from Iranian Persian, and human preference data.1

Validated read-speech hours in Mozilla Common Voice 25.0
Pashto (ps)3,041.86 h
Persian (fa)372.94 h
Common Voice Scripted Speech is read-aloud speech by design. It is not conversation, and it is not preference data. Sources 12
3,041.86hoursValidated Pashto read speech in Mozilla Common Voice 25.0 (March 2026), from 8,239 speakers1
372.94hoursValidated Persian read speech in the same release, from 4,655 speakers2
6.4percentShare of Persian clips in that release from speakers who declared as female2
60–80millionEstimated Pashto speakers across Afghanistan, Pakistan and the diaspora3
8consonantsPashto consonants absent from Arabic and Persian keyboards3

Conversation, not read-aloud sentences

Scripted read speech teaches a model to recognise careful reading. Assistants hear interruptions, code-switching, background noise and regional pronunciation, and need recordings of real conversation to learn them.

Clean Pashto script

Eight Pashto consonants are missing from Arabic and Persian keyboards, and informal text routinely drops or substitutes them. Data collected without a normalization standard teaches a model to misspell Pashto.3

Afghan Dari, not Iranian Persian

Dari and Iranian Persian share a script and most of their grammar but differ in everyday vocabulary, idiom and pronunciation. Benchmarks such as FLORES-200 list them as separate languages; many training sets still do not.4

Dialect and gender balance

Volunteer corpora follow whoever volunteers. In Common Voice’s Persian release, 6.4 percent of clips come from speakers who declared as female. Balanced speaker panels have to be specified and recruited, not assumed.2

Human preference data

Preference datasets are expensive to build, and the widely used ones are dominated by English and a few large languages. Pashto and Dari reward models need judgments from native raters who read the dialect, the register and the culture.

Where Pashto and Dari training data goes wrong

Most failures in Afghan-language data are invisible to anyone who does not read the language letter by letter. These are the ones that most often decide whether a dataset improves a model or teaches it mistakes.

Eight Pashto letters that break training data

Pashto adds eight consonants to the Arabic script. When a keyboard lacks them, writers fall back on a nearby letter, usually the one that matches their own pronunciation, and the word changes. Every file we deliver is checked letter by letter against a written standard.3
The eight Pashto consonants and the letters they are often typed as
Pashto letterCode pointOften typed asCode point
Pashto letterټCode pointU+067COften typed asتCode pointU+062A
Pashto letterډCode pointU+0689Often typed asدCode pointU+062F
Pashto letterړCode pointU+0693Often typed asرCode pointU+0631
Pashto letterڼCode pointU+06BCOften typed asنCode pointU+0646
Pashto letterښCode pointU+069AOften typed asش or خCode pointU+0634, U+062E
Pashto letterږCode pointU+0696Often typed asژ or گCode pointU+0698, U+06AF
Pashto letterځCode pointU+0681Often typed asزCode pointU+0632
Pashto letterڅCode pointU+0685Often typed asسCode pointU+0633

Why Iranian Persian data cannot stand in for Dari

The same sentence written for an Afghan reader and for an Iranian reader uses different everyday words. A model trained on Iranian Persian answers Afghan users in a register they recognise as foreign.
Everyday words that differ between Afghan Dari and Iranian Persian
EnglishAfghan DariIranian Persian
EnglishHospitalAfghan DariشفاخانهshafākhānaIranian Persianبیمارستانbimārestān
EnglishCarAfghan DariموترmotarIranian Persianماشینmāshin
EnglishUniversityAfghan DariپوهنتونpohantunIranian Persianدانشگاهdāneshgāh
EnglishProvinceAfghan DariولایتwelāyatIranian Persianاستانostān
EnglishNational ID cardAfghan DariتذکرهtazkiraIranian Persianکارت ملیkārt-e melli
EnglishTomatoAfghan Dariبادنجان رومیbādenjān-e rumiIranian Persianگوجه‌فرنگیgojeh-farangi
Illustrative pairs. Program vocabularies are built and reviewed by native Dari linguists for each domain.

One normalization rule does not fit both languages

Pipelines built for Persian merge letters that Pashto keeps apart. The same cleaning script that fixes Dari text corrupts Pashto.

Pashto

  • ي (U+064A) and ی (U+06CC) are different Pashto letters with different sounds, and are never merged
  • ې (U+06D0), ۍ (U+06CD) and ئ (U+0626) are kept as separate letters
  • Pashto ګ (U+06AB) is kept distinct from Persian گ (U+06AF)
  • The eight Pashto consonants are restored wherever a substitute was typed

Dari

  • Arabic Yeh ي (U+064A) becomes Farsi Yeh ی (U+06CC)
  • Arabic Kaf ك (U+0643) becomes Keheh ک (U+06A9)
  • The zero-width non-joiner (U+200C) is kept inside compounds such as می‌روم
  • Digits follow the specification: ۰۹ (U+06F0–U+06F9) or Western digits

What a native preference rater catches

Preference data is only as good as the rater’s reading. In each example below the two responses look equivalent to a rater who does not read the dialect, and they are not.

Pashto: the letters a keyboard drops

Prompt
Translate into Pashto: “Please check the car first.”
Response A
مهربانی وکرئ لومری موتر وگورئ.
Typed on a Persian keyboard: five letters substituted.
Response B Preferred
مهرباني وکړئ لومړی موټر وګورئ.
Written in correct Pashto script.
Rubric for the Pashto example
CriterionResponse AResponse B
CriterionMeaningResponse ACorrectResponse BCorrect
CriterionScript fidelityResponse Aر for ړ twice, ت for ټ, گ for ګ, ی for يResponse BCorrect Pashto letters
CriterionTraining-data riskResponse ATeaches the model keyboard substitutionsResponse BClean
Rater’s rationaleAt a glance the two responses look identical. A rater who does not read Pashto letter by letter marks them equal, and the reward model learns to accept the substitutions.

Dari: the word a user expects

Prompt
موترم روشن نمی‌شود. اول چه چیز را چک کنم؟
My car won’t start. What should I check first?
Response A
اول باتری ماشین را چک کنید.
First, check the car battery. (Uses ماشین, the Iranian Persian word for car.)
Response B Preferred
اول باتری موتر را چک کنید.
First, check the car battery. (Uses موتر, the Afghan Dari word the user wrote.)
Rubric for the Dari example
CriterionResponse AResponse B
CriterionMeaningResponse ACorrectResponse BCorrect
CriterionDialect fitResponse AIranian Persian vocabularyResponse BAfghan Dari, as the user wrote
CriterionRegisterResponse AReads as foreign to the userResponse BNatural
Rater’s rationaleBoth answers give the same advice. B keeps the user’s own word for car; A switches to the Iranian Persian word, which an Afghan user reads as a foreign register. A rater without native Dari judgment is likely to mark them equal.

Spoken: the accent a listener hears

Prompt
A speaker from Kandahar asks a voice assistant, in Southern Pashto, where the nearest pharmacy is.
Response A
The reply is grammatical, but pronounces ښ as speakers of Northern Pashto do.
Response B Preferred
The reply is grammatical and pronounces ښ as the speaker does.
Rubric for the Spoken example
CriterionResponse AResponse B
CriterionMeaningResponse ACorrectResponse BCorrect
CriterionDialect fitResponse ANorthern pronunciationResponse BSouthern pronunciation, as asked
CriterionNaturalnessResponse ASounds like another regionResponse BSounds local
Rater’s rationaleOnly a native listener hears the difference. Spoken-response preference data captures pronunciation, dialect fit and naturalness, the dimensions a text-only rater cannot judge.

Pashto and Dari dialects, and the other Afghan languages

Every specification names the varieties, regions and speaker balance a program needs, and every datasheet reports what was delivered against it.

Pashto varieties

  • Southern Pashto @@PROTECT0@@Kandahar and the south; the language is called Pashto.
  • Northern Pashto @@PROTECT1@@Including the speech of Peshawar, where the language is called Pakhto.
  • Central Pashto @@PROTECT2@@The varieties between the two, with their own sounds for ښ and ږ.
The letter ښ alone is pronounced differently across these varieties, which is why the same language is called Pashto in Kandahar and Pakhto in Peshawar.

Dari varieties

  • Kabuli Dari @@PROTECT3@@The broadcast and administrative standard.
  • Herati Dari @@PROTECT4@@Western Afghanistan, with its own vocabulary and intonation.
  • Dari of the north and north-east @@PROTECT5@@Including the speech of Badakhshan and the northern provinces.
  • Hazaragi @@PROTECT6@@A variety of Dari, treated as operationally separate.

Program depth across the 24 Afghan languages

Production
Full programs in text, speech and preference data.
Scoped
Programs scoped to the speakers, writing system and volume available.
Feasibility first
A written feasibility study before any volume is promised.
Program depth across the 24 Afghan languages: family, ISO 639-3 code and program depth
LanguageFamilyISO 639-3Program depth
LanguagePashtoFamilyIranianISO 639-3pbt, pbu, pstProgram depthProduction
LanguageDariFamilyIranianISO 639-3prsProgram depthProduction
LanguageHazaragi (a variety of Dari)FamilyIranianISO 639-3hazProgram depthScoped
LanguageUzbekiFamilyTurkicISO 639-3uzsProgram depthScoped
LanguageTurkmeniFamilyTurkicISO 639-3tukProgram depthScoped
LanguageBalochiFamilyIranianISO 639-3bgnProgram depthScoped
LanguagePashayiFamilyIndo-AryanISO 639-3aee, glh, psh, psiProgram depthScoped
LanguageAimaqFamilyIranianISO 639-3aiqProgram depthScoped
LanguageKyrgyzFamilyTurkicISO 639-3kirProgram depthFeasibility first
LanguageOrmuriFamilyIranianISO 639-3oruProgram depthFeasibility first
LanguageParachiFamilyIranianISO 639-3prcProgram depthFeasibility first
LanguageWakhiFamilyIranian (Pamir)ISO 639-3wblProgram depthFeasibility first
LanguageShughniFamilyIranian (Pamir)ISO 639-3sghProgram depthFeasibility first
LanguageSanglechiFamilyIranian (Pamir)ISO 639-3sgyProgram depthFeasibility first
LanguageIshkashimiFamilyIranian (Pamir)ISO 639-3iskProgram depthFeasibility first
LanguageMunjiFamilyIranian (Pamir)ISO 639-3mnjProgram depthFeasibility first
LanguageYidghaFamilyIranian (Pamir)ISO 639-3ydgProgram depthFeasibility first
LanguageGawarbatiFamilyIndo-AryanISO 639-3gwtProgram depthFeasibility first
LanguageTirahiFamilyIndo-AryanISO 639-3traProgram depthFeasibility first
LanguageNuristani (Ashkun group)FamilyNuristaniISO 639-3askProgram depthFeasibility first
LanguageKatiFamilyNuristaniISO 639-3bshProgram depthFeasibility first
LanguagePrasunFamilyNuristaniISO 639-3prnProgram depthFeasibility first
LanguageWaigaliFamilyNuristaniISO 639-3wbkProgram depthFeasibility first
LanguageBrahuiFamilyDravidianISO 639-3brhProgram depthFeasibility first

How we deliver a Pashto and Dari data program

Seven steps, each ending in a document you can check. Nothing moves to the next step until the one before it has passed.

  1. 01

    Data specification

    We define the target model, data mix, dialects, speaker panel, domains, volume, formats and acceptance thresholds in a written specification you approve.
    Output: Signed data specification
  2. 02

    Guidelines and gold set

    Native linguists write annotation guidelines with Pashto and Dari examples and build a gold set that every annotator must pass.
    Output: Guidelines and gold set
  3. 03

    Recruitment and consent

    Contributors are recruited through our own Afghan diaspora networks, consented in their own language and paid directly by Ariana Nexus.
    Output: Consent ledger
  4. 04

    Paid pilot

    A small batch runs end to end at full fidelity, is scored against the thresholds and is reviewed with your team before any scale-up.
    Output: Pilot batch and scorecard
  5. 05

    Production

    Items that need judgment are annotated twice, with gold items embedded throughout and calibration sessions whenever agreement drifts.
    Output: Versioned batches
  6. 06

    Adjudication and quality assurance

    A senior linguist resolves disagreements, and agreement, error and audit statistics are reported with every batch.
    Output: Batch QA report
  7. 07

    Delivery and documentation

    Files go to your storage with a datasheet, the final guidelines, the QA report and the provenance record.
    Output: Release with datasheet
  8. Delivery formats

    Text and preference data
    JSONL, Parquet, CSV, TSV or the Hugging Face Datasets layout
    Audio
    WAV or FLAC at 8, 16 or 48 kHz, 16- or 24-bit
    Alignment
    Praat TextGrid, ELAN EAF and RTTM
    Subtitles
    SRT and WebVTT
    Delivery
    Amazon S3, Google Cloud Storage, Azure Blob Storage or SFTP
    Every release ships with a datasheet, the final guidelines, a QA report and the provenance and consent record.

How Pashto and Dari annotation quality is measured

Thresholds are written into the specification before work starts, and every batch reports against them. The defaults below are where most programs begin.5

Quality measures and default thresholds
MeasureWhat it tells youDefault threshold, set per program
MeasureInter-annotator agreement (Krippendorff’s alpha or Cohen’s kappa)What it tells youWhether two trained annotators independently reach the same label.Default threshold, set per programAlpha of 0.80 or higher on categorical labels, reported per label
MeasureGold-item accuracyWhat it tells youWhether each annotator follows the guideline, measured on hidden items with known answers.Default threshold, set per program95 percent or higher per annotator, per batch
MeasureTranscription accuracy auditWhat it tells youWord and character error rate against an expert re-transcription of a random sample.Default threshold, set per program5 percent word error or lower on clean speech; character error reported
MeasureScript conformanceWhat it tells youWhether every file follows the Pashto or Dari normalization standard.Default threshold, set per programEvery file, checked automatically
MeasureRater consistencyWhat it tells youPosition-bias checks and repeated items that show whether preference raters judge consistently.Default threshold, set per programRaters below threshold are recalibrated or removed
MeasureBatch acceptance samplingWhat it tells youWhether a batch meets the agreed quality level before release, using ISO 2859-1 sampling plans.Default threshold, set per programAcceptance quality limit agreed in the specification
MeasureConsent coverageWhat it tells youWhether every recording and text traces to a valid consent record.Default threshold, set per programEvery item, verified before delivery

Why Ariana Nexus for Pashto and Dari training data

Afghan-language data is judged by people who can read what the model got wrong. That is the difference between a dataset that improves a model and one that teaches it mistakes.

Scholars, not a crowd

Our linguists and annotators are university-educated native speakers of Pashto and Dari, led by alumni and scholars of Cornell, the University of Chicago and the University of British Columbia. No certification exists for Afghan-language annotation, so we set the standard ourselves: written guidelines, gold sets and calibration that every annotator passes before touching client data.

Dialect control, in writing

Every specification names the varieties, regions and speaker balance required, and every datasheet shows what was delivered against it.

Community trust

Afghan-led recruitment through our own diaspora networks, consent in the contributor’s language and pay agreed before the first recording.

A second review for culture and faith

Our Cultural Compliance Bureau reviews content touching religion, gender, ethnicity and the war, so sensitive data is labeled with context rather than guesswork.

In-house, end to end

Every part is produced by our own people: one engagement, one point of accountability, no subcontracted vendors.

Evidence you can audit

Datasheets, agreement statistics, consent ledgers and provenance fields ready for EU AI Act and California AB 2013 disclosures.

The team behind Ariana Nexus training data

Pashto and Dari data programs are led by the firm’s partners and principals, alumni and scholars of Cornell University, the University of Chicago and the University of British Columbia, and delivered by a bench of university-educated native linguists.

Each program has one accountable lead, a linguistic lead for each language and a measurement lead who owns the numbers. The people who write the guidelines also adjudicate the hardest items, so the standard and the judgment stay in the same hands.
Program oversight

Hassan Ukasha

Managing Partner, Washington, D.C.
  • B.S., Cornell University
  • M.P.H., Cornell University

Oversees the firm’s operations and serves as executive sponsor for every Pashto and Dari data program: scope, contributor safety, data rights and final sign-off before any dataset leaves the firm.

Languages
Pashto, Dari, English, Urdu, Hindi; working Arabic
Zeba Haqbani

Zeba Haqbani

Senior Partner
  • B.Sc., University of British Columbia
Data platform and delivery: secure annotation workspaces, versioning and file delivery.
Hussain Ahmad

Hussain Ahmad

Principal
  • M.Eng., Cornell University
  • Ph.D., University of Chicago
Dataset design and measurement: sampling plans, agreement statistics and datasheets.
Wasil Peroz

Wasil Peroz

Principal
  • B.A., Milli University
  • M.Sc., Otto von Guericke University Magdeburg
Consent, licensing and data rights: contributor agreements, GDPR and provenance records.
Maryam Safi

Maryam Safi

Principal
  • B.A., Cornell University
Linguistic program lead: guidelines, annotator calibration and adjudication.

The linguist bench

Behind them is a bench of native Pashto and Dari linguists and annotators: university graduates and current scholars from the Afghan diaspora, matched to each program’s dialects and domains. Bench members are not named publicly, to protect those whose families remain in Afghanistan.
An empty conference room with a long table and windows.

Where Pashto and Dari training data is used

LLM post-training

Supervised fine-tuning, RLHF and DPO for assistants that answer in Pashto and Dari.

Voice assistants and speech-to-speech models

Spoken prompts, conversation and spoken-response preference for voice products.

Machine translation

Parallel corpora and translation preference for Pashto, Dari and English.

Trust and safety classifiers

Labeled hate speech, harassment, extremism and misinformation in Pashto and Dari.

Search and retrieval

Queries, relevance judgments and entity labels for Afghan-language search.

Document AI and OCR

Ground truth for printed and handwritten Pashto and Dari.

Healthcare AI

Clinical conversation and patient-facing language, labeled by health-literate annotators.

Public-sector and humanitarian services

Information services for Afghan communities in the United States, Europe and beyond.

How Pashto and Dari data engagements are structured

Paid pilot

A full-fidelity sample, typically delivered two to four weeks after the specification is signed, scored against your thresholds.

Fixed-scope dataset

A defined dataset delivered in versioned batches, each with a QA report and a datasheet.

Dedicated annotation team

A named, calibrated team working in your annotation platform or in our secure workspace.

Continuous preference program

Recurring preference collection for iterative RLHF cycles, with a stable, calibrated rater pool.
Programs are priced per unit (audio hour, item or preference pair) or per dedicated team, once the pilot has measured real throughput.

Key terms in Pashto and Dari training data

Supervised fine-tuning (SFT) data
Prompts paired with ideal responses, used to teach a model a task or a style.
Preference data
Human judgments of which of two or more model responses is better, and why.
RLHF
Reinforcement learning from human feedback: training a model against a reward model learned from preference data.
DPO
Direct preference optimization: training a model directly on chosen and rejected response pairs, without a separate reward model.
Inter-annotator agreement
A statistic, such as Krippendorff’s alpha, showing how often trained annotators independently assign the same label.
Gold set
Items with verified answers, hidden among real work to measure each annotator’s accuracy.
Datasheet
A document describing how a dataset was made: sources, contributors, consent, annotation, known limits and intended use.
Pashto (ps)
One of Afghanistan’s two official languages, spoken across Afghanistan and Pakistan and called Pakhto in the north and east; ISO 639-1 code ps.
Dari (prs)
Afghanistan’s other official language and the Afghan variety of Persian; ISO 639-3 code prs, distinct from Iranian Persian.
Zero-width non-joiner
An invisible character (U+200C) that separates parts of a Dari word without a space; losing it changes how the text is tokenized.

Pashto and Dari training data: frequently asked questions

What is Pashto and Dari AI training data?

It is the text, speech and human feedback an AI model learns from in Afghanistan’s two main languages: instructions and ideal answers, conversations, domain text, parallel translations, recorded speech with transcripts, and preference judgments comparing model responses. Ariana Nexus creates, records and labels this data with native Afghan linguists and documents every dataset.

Which company provides Pashto and Dari training data and annotation?

Ariana Nexus, a consulting and professional services firm headquartered in Washington, D.C., collects and annotates Pashto and Dari text, speech and preference data for AI labs, technology companies, government contractors and researchers. Programs are led by native Afghan linguists and scholars and delivered in-house, without subcontracted vendors.

Can Iranian Persian (Farsi) data be used to train a Dari model?

Only as a starting point. Dari and Iranian Persian share a script and most of their grammar, but differ in everyday vocabulary, idiom and pronunciation: hospital is shafakhana in Dari and bimarestan in Iran. A model trained mainly on Iranian Persian answers Afghan users in a register they recognise as foreign. Dari needs Dari data and Dari raters.

Which Pashto and Dari dialects do you cover?

For Pashto: Southern, Central and Northern varieties, including Pakhto as spoken in Peshawar. For Dari: Kabuli, Herati and the Dari of the north and north-east, plus Hazaragi, a variety of Dari. Every specification names the varieties and speaker balance required, and the datasheet reports what was delivered.

What is RLHF preference data, and do you collect it in Pashto and Dari?

Preference data records which of two or more model responses a person prefers, and why. It trains the reward models behind RLHF and the chosen and rejected pairs behind DPO. We collect pairwise judgments, rankings, rubric ratings, rewrites and safety preferences in Pashto and Dari from native raters, with written rationales.

How do you measure annotation quality?

Every program sets written thresholds before work starts. We report inter-annotator agreement (Krippendorff’s alpha or Cohen’s kappa), gold-item accuracy for each annotator, transcription error rates from expert re-transcription, script conformance for every file and rater consistency for preference data. Disagreements are adjudicated by a senior linguist.

How do you handle consent and privacy for Pashto and Dari speech data?

Contributors give informed consent in Pashto or Dari, read aloud where needed, with the right to withdraw, and they are paid directly. Identities are stored apart from recordings, transcripts are redacted, and no data or payment passes through channels controlled by the de facto authorities in Afghanistan.

Can you annotate sensitive content, including material from the war in Afghanistan?

Yes. Trust and safety programs often need violent, extremist or conflict-related Pashto and Dari content labeled accurately. Our annotators work with capped exposure and scheduled debriefing, and our Cultural Compliance Bureau reviews items touching religion, ethnicity and the war, so labels reflect context rather than keywords.

What formats do you deliver Pashto and Dari datasets in?

Text and preference data in JSONL, Parquet, CSV or the Hugging Face Datasets layout; audio in WAV or FLAC at the sample rate your model needs; alignments in Praat TextGrid, ELAN or RTTM; subtitles in SRT or WebVTT. Every release includes a datasheet, the final guidelines and a QA report.

How long does a Pashto or Dari data program take?

A paid pilot is typically delivered two to four weeks after the specification is signed. Production timelines depend on volume, dialect mix and recording conditions, and are set from the pilot’s measured throughput rather than estimated in advance.

Do you collect data in other Afghan languages?

Yes, by program. Pashto and Dari run at production scale. Hazaragi, Uzbeki, Turkmeni, Balochi, Pashayi and Aimaq are scoped to the speakers and writing systems available. The remaining Afghan languages begin with a written feasibility study before any volume is promised.

How are Pashto and Dari data programs priced?

Per unit (audio hour, annotated item or preference pair) or per dedicated team. Every program starts with a paid pilot, and production pricing is set from what the pilot measures: throughput, agreement and rework. We do not price by the piece to anonymous crowds.

Can Ariana Nexus support U.S. federal agencies and prime contractors?

Yes. Ariana Nexus is a U.S. firm headquartered in Washington, D.C. and registered in SAM.gov (UEI M2UDMUDFXGL9, CAGE 1Z3A3). For federal programs, contributors, annotators and data handling can be kept inside the United States when the program requires it.

Scope a Pashto and Dari data program

Tell us the model, the languages and the data you need. We will return a written data specification and a pilot plan.

Request a pilot scope

Sources

  1. Mozilla Data Collective. Common Voice Scripted Speech 25.0, Pashto. Datasheet for cv-corpus-25.0-2026-03-09. https://mozilladatacollective.com/datasets/cmndf6mgs001lnz07bf9t3skp
  2. Mozilla Data Collective. Common Voice Scripted Speech 25.0, Persian. Datasheet for cv-corpus-25.0-2026-03-09. https://datacollective.mozillafoundation.org/datasets/cmn2gho8i01gio107ckfuqzxo
  3. Pashto Common Voice: Building the First Open Speech Corpus for a 60-Million-Speaker Low-Resource Language. arXiv:2603.27021, 2026. https://arxiv.org/abs/2603.27021
  4. NLLB Team. FLORES-200 evaluation benchmark, language codes prs_Arab (Dari), pbt_Arab (Southern Pashto) and pes_Arab (Western Persian). https://github.com/facebookresearch/flores/blob/main/flores200/README.md
  5. ISO/IEC 5259-4:2024. Artificial intelligence: Data quality for analytics and machine learning, Part 4: Data quality process framework. https://www.iso.org/standard/81093.html
  6. California Assembly Bill 2013 (2024). Generative artificial intelligence: training data transparency. https://leginfo.legislature.ca.gov/faces/billNavClient.xhtml?bill_id=202320240AB2013
  7. Regulation (EU) 2024/1689, the Artificial Intelligence Act, Articles 10 and 53. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
  8. Gebru, T. et al. Datasheets for Datasets. Communications of the ACM, 2021. https://arxiv.org/abs/1803.09010