Afghan Language Training Data — Pashto and Dari Annotation, RLHF, and Evaluation Sets
Ariana Nexus builds Afghan language training data for AI teams — annotation, RLHF and preference data, and dialect-aware evaluation sets across 24 Afghan languages, Pashto and Dari first. Every dataset is produced by Afghan native speakers, validated for culture as well as language, and delivered with provenance an auditor can follow.
Ariana Nexus is a Washington, D.C.–area firm providing Afghan language services and cultural intelligence — interpretation, translation, cultural training, compliance support, and AI data — across 24 Afghan languages. Never extractive.
Exhibit 01 — The coverage gap
Yong, Menghini & Bach, 2023 (arXiv:2310.02446).
What makes Afghan language training data different?
Afghan language training data is text, speech, and preference data built for Pashto, Dari, and the other 22 Afghan languages by people who speak them. It differs from general multilingual data in three ways: the dialects are held apart rather than flattened, Afghanistan Dari is never substituted with Iranian Persian, and cultural judgment is recorded alongside the label.
RLHF, or reinforcement learning from human feedback, is model training on the rankings people give to competing answers. A gold set is a small expert-built reference sample that every annotator is measured against. Inter-annotator agreement is the rate at which independent annotators give the same label to the same item.
Independent testing of the models this data trains is a separate Ariana Nexus capability: Afghan language LLM evaluation and red-teaming
Cultural and religious defects in model output are classified and severity-scored under Afghan cultural red-teaming for AI models
You cannot align your way out of bad data.
A model is only as good as the data it learns from, and for Pashto and Dari that data is the problem. Most of what was available to train on was web-scraped, machine-translated, dialect-flattened, and culturally unvalidated, and the model inherited every flaw in it. No volume of downstream RLHF repairs a foundation that erased the dialects and misread the culture; it only makes the errors more fluent.
The fix is upstream, and it is human. Good alignment needs good human feedback — expert, costly, and for Afghan languages scarce. It has to come from Afghan native speakers who know the dialects, judge the culture, and produce data that holds up to inter-annotator agreement and gold-standard validation.
That work also has a human cost. Independent research finds the major data-labor platforms scoring at or below the bare minimum on fair pay and conditions, with documented worker harm. Data made cheaply is often made extractively, and that is a liability of its own — in the supply chain and in the conscience.
Ariana Nexus is the Data Factory built the other way: dialect-aware Pashto and Dari data, and data in the other 22 Afghan languages, for training, fine-tuning, RLHF, and evaluation — produced by Afghan native speakers, traceable to its source, and made by people paid and treated well.
of practitioners' time goes to data preparation, not model architecture (Anaconda, 2020)
the score independent research gives major data-labor platforms on fair pay
the Ariana Nexus dialect-aware gold-set method
One practice. Three coordinated capabilities — anchored in the Factory.
Three institutional capabilities, joined into data a model can actually learn from.
ADF · AI Data Factory
The data-production pipeline: supervised fine-tuning sets, RLHF and preference data, instruction tuning, and dialect-aware evaluation references. Pashto data annotation and Dari text annotation run to written guidelines, with inter-annotator agreement, gold-standard checks, and provenance and lineage recorded on every dataset.
Protocol: The ADF Pipeline
HIC · Human Intelligence Collective
Afghan native-speaker annotators, dialect experts, and human-feedback providers across all 24 Afghan languages — Pashto and Dari first — matched to the dialect the work requires. This is the human judgment that makes the data real, treated fairly and protected, never extractive.
Protocol: The Dialect Reference Standard
CCB · Cultural Compliance Bureau
Cultural validation of the data itself, not just the language: annotation-guideline review, worker-wellbeing and data-ethics governance, provenance and consent audit, and the CCB Sign-Off Mark on every delivered dataset.
Protocol: The CCB Sign-Off Mark
How Ariana Nexus produces the data: the ADF Pipeline
Five pipeline stages, three capabilities, five validation gates. The ADF Pipeline produces the data; the Five-Gate Validation Protocol governs its quality, culture, provenance, and the conditions under which it was made.
The five stages
Ingest & Scope
Model, task, languages, dialects
Annotate · HIC
Native-speaker labeling & judgment
Preference · RLHF
Human feedback & ranking
Five-Gate QA
Validation & CCB sign-off
Governed Delivery
Data cards & provenance
The Five Gates
Linguistic Accuracy — data linguistically accurate across 24 languages and dialects; inter-annotator agreement and gold-standard validation; the Pashto-Dari Parity Index applied.
Cultural Validity — data culturally valid, not merely literally correct; dialect-aware; cleared by the CCB Sign-Off Mark.
Standards Conformance — data-quality and annotation standards (ISO/IEC 5259, ISO/IEC 25012); RLHF and preference-data best practice; data provenance, consent, and licensing, with EU AI Act training-data transparency where applicable.
Population Risk — ethical labor: fair, living compensation, dignity, and never extractive; worker-wellbeing protection for sensitive or distressing content; consent and privacy; no rights-infringing data sourcing.
Institutional Sign-Off — datasets documented with provenance, lineage, quality metrics, and data cards — traceable and audit-ready.
The Four-Phase Delivery Cycle
Situation — Understand
The model, the task (SFT, RLHF, evaluation), the target languages and dialects, and the data requirements mapped.
Complication — Architect
The ADF Pipeline configured; annotation guidelines, dialect-aware reference design, and the quality framework built; the annotator cohort assembled ethically.
Resolution — Deploy
Data produced — annotated, human feedback collected, dialect-aware reference sets built — with QA and provenance throughout.
Measured Outcome — Govern
Data quality measured (agreement, gold-standard, downstream evaluation); data cards and provenance delivered; pipelines maintained and refreshed.
Standards & compliance
Mapped to the registries an ML-data lead, a procurement officer, and a data-governance reviewer recognize.
Data Quality & ML Data
ISO/IEC 25012ISO/IEC 5259ISO/IEC 8183Inter-annotator agreement & gold-standard practiceAI & Provenance
NIST AI RMF — data governanceEU AI Act — training-data transparencyData provenance & lineage standardsRLHF & preference-data best practiceWhat governs AI training data — and where it stands.
The binding and voluntary regimes that reach training data, RLHF, and data labor in a healthcare context. Status as of June 2026.
Voluntary frameworks (GMLP, PCCP principles, NIST AI RMF, ISO/IEC) are consensus standards, not law; HIPAA, Section 1557, ONC HTI-1, and the EU AI Act bind on their stated timelines. EU AI Act high-risk data-governance duties apply 2 August 2026 (the Act entered into force 1 August 2024); verify status at publish.
Healthcare
Where the data meets the clinic.
For health systems and their compliance officers, the training-data layer is now a regulated surface — a matter of civil rights, certification, and patient safety.
Section 1557 · in force 1 May 2025
Decision-support tools are in scope of civil-rights law.
The 2024 Section 1557 rule requires covered entities to identify and mitigate discrimination risk in patient-care decision-support tools — a duty that runs upstream into the data those tools learn from.
ONC HTI-1
Certification now demands training-data transparency.
HTI-1 requires certified health IT to surface 31 source attributes for predictive tools, so clinicians can judge whether a model is Fair, Appropriate, Valid, Effective, and Safe (FAVES) — including how it was trained.
Obermeyer, Science 2019
Bias is set in the label, not just the model.
The canonical clinical-AI failure traces to a cost-as-proxy training target, not architecture — proof that representative, well-specified data is itself a patient-safety control, not a downstream nicety.
HIPAA §164.514
PHI in training corpora is a hard prerequisite.
De-identification — Safe Harbor's 18 identifiers or documented Expert Determination — governs any PHI flowing into training data; unstructured clinical notes are the weak point, retaining residual identifiers.
What partnership looks like
Foundations
Scoped and mapped. The model, the task, the Afghan languages and dialects in scope, and the volume of data the work will need.
Activation
Built to standard. The pipeline configured, annotation guidelines written, and the Afghan annotator cohort assembled and paid fairly.
Operating Rhythm
The active state. Data produced with QA and provenance; quality measured against agreement and gold standards.
Continuous Stewardship
Refreshed and scaled. Pipelines maintained; data cards and provenance kept current as the model and task evolve.
Leadership
Who leads the AI & Data Systems Practice
Naseer Ahmadzai
Senior Partner, AI Data Factory
Data-production, ML-data, and annotation-operations leadership.
Wahida Noori
Director, Annotation & Human Feedback (RLHF)
RLHF, preference-data, and annotation-platform leadership.
Sahar Rahimi
Director, Data Ethics & Provenance
Responsible-data-sourcing, worker-wellbeing, and data-governance leadership; owns the ethical-annotation and provenance commitments.
Questions AI teams ask before they buy
What is Afghan language training data?
Afghan language training data is the labeled text, audio, and human-preference data used to train, fine-tune, and evaluate models in Afghan languages. Ariana Nexus produces Pashto training data, Dari training data, and sets in the other 22 Afghan languages — annotation, RLHF and preference data, instruction tuning, and dialect-aware evaluation references, each delivered with a data card.
Can a Farsi or Persian dataset be used for Dari?
Not safely. Afghanistan Dari and Iranian Persian differ in vocabulary, register, and idiom, and a model trained on Persian data produces Dari that Afghan users read as foreign or simply wrong. Ariana Nexus builds Afghanistan Dari data with Afghan native speakers, and keeps Kabuli, Herati, Badakhshani, and Hazaragi — a dialect of Dari — labeled separately rather than merged.
How much does Afghan language training data cost?
Cost depends on the language, the dialect coverage, the task, and the volume. An evaluation reference set of a few thousand items is a different order of work from a full RLHF preference program in Pashto and Dari. Ariana Nexus scopes each program against the model and the task before quoting, and publishes no rate card, because a rate quoted without scope is a guess.
Is Hazaragi a separate language or a Dari dialect?
Hazaragi is a dialect of Dari, not a separate language, although many buyers search for it by name. It carries distinct vocabulary and pronunciation, and an Afghan language dataset that merges it into standard Dari loses exactly the variation that matters. Ariana Nexus labels Hazaragi as a Dari variety and holds it separate in the data.
Should annotators be described as Pashto or Pashtun speakers?
Pashto is the language; Pashtun is the people who speak it, so the correct phrase is a Pashto annotator or a Pashto speaker. The same distinction applies to Afghan and Afghani: Afghan describes the people and the languages, while the afghani is Afghanistan's currency. Ariana Nexus uses the language names in every data card and annotation guideline.
How are Afghan annotators recruited and paid?
Ariana Nexus recruits Afghan native speakers from a graduate-level bench, matched to the dialect the work requires, and pays them fairly with worker-wellbeing protection for sensitive or distressing content. Consent, licensing, and annotation guidelines are documented for every dataset, so the provenance a regulator asks for is already on file.
Request an AI Data Factory Consultation.
Request a confidential briefingEngagements are governed by the Ariana Nexus Trust Center.