Afghan Language Training Data — Pashto and Dari Annotation, RLHF, and Evaluation Sets

Ariana Nexus builds Afghan language training data for AI teams — annotation, RLHF and preference data, and dialect-aware evaluation sets across 24 Afghan languages, Pashto and Dari first. Every dataset is produced by Afghan native speakers, validated for culture as well as language, and delivered with provenance an auditor can follow.

Ariana Nexus is a Washington, D.C.–area firm providing Afghan language services and cultural intelligence — interpretation, translation, cultural training, compliance support, and AI data — across 24 Afghan languages. Never extractive.

Exhibit 01 — The coverage gap

High-resource coverageAfghan-language data — sparse
~79%of unsafe prompts in low-resource languages bypassed GPT-4 safety — versus under 1% in English.

Yong, Menghini & Bach, 2023 (arXiv:2310.02446).

What makes Afghan language training data different?

Afghan language training data is text, speech, and preference data built for Pashto, Dari, and the other 22 Afghan languages by people who speak them. It differs from general multilingual data in three ways: the dialects are held apart rather than flattened, Afghanistan Dari is never substituted with Iranian Persian, and cultural judgment is recorded alongside the label.

RLHF, or reinforcement learning from human feedback, is model training on the rankings people give to competing answers. A gold set is a small expert-built reference sample that every annotator is measured against. Inter-annotator agreement is the rate at which independent annotators give the same label to the same item.

Independent testing of the models this data trains is a separate Ariana Nexus capability: Afghan language LLM evaluation and red-teaming

Cultural and religious defects in model output are classified and severity-scored under Afghan cultural red-teaming for AI models

You cannot align your way out of bad data.

A model is only as good as the data it learns from, and for Pashto and Dari that data is the problem. Most of what was available to train on was web-scraped, machine-translated, dialect-flattened, and culturally unvalidated, and the model inherited every flaw in it. No volume of downstream RLHF repairs a foundation that erased the dialects and misread the culture; it only makes the errors more fluent.

The fix is upstream, and it is human. Good alignment needs good human feedback — expert, costly, and for Afghan languages scarce. It has to come from Afghan native speakers who know the dialects, judge the culture, and produce data that holds up to inter-annotator agreement and gold-standard validation.

That work also has a human cost. Independent research finds the major data-labor platforms scoring at or below the bare minimum on fair pay and conditions, with documented worker harm. Data made cheaply is often made extractively, and that is a liability of its own — in the supply chain and in the conscience.

Ariana Nexus is the Data Factory built the other way: dialect-aware Pashto and Dari data, and data in the other 22 Afghan languages, for training, fine-tuning, RLHF, and evaluation — produced by Afghan native speakers, traceable to its source, and made by people paid and treated well.

~45%

of practitioners' time goes to data preparation, not model architecture (Anaconda, 2020)

Bare minimum

the score independent research gives major data-labor platforms on fair pay

The Dialect Reference Standard

the Ariana Nexus dialect-aware gold-set method

One practice. Three coordinated capabilities — anchored in the Factory.

Three institutional capabilities, joined into data a model can actually learn from.

ADF · AI Data Factory

The data-production pipeline: supervised fine-tuning sets, RLHF and preference data, instruction tuning, and dialect-aware evaluation references. Pashto data annotation and Dari text annotation run to written guidelines, with inter-annotator agreement, gold-standard checks, and provenance and lineage recorded on every dataset.

Protocol: The ADF Pipeline

HIC · Human Intelligence Collective

Afghan native-speaker annotators, dialect experts, and human-feedback providers across all 24 Afghan languages — Pashto and Dari first — matched to the dialect the work requires. This is the human judgment that makes the data real, treated fairly and protected, never extractive.

Protocol: The Dialect Reference Standard

CCB · Cultural Compliance Bureau

Cultural validation of the data itself, not just the language: annotation-guideline review, worker-wellbeing and data-ethics governance, provenance and consent audit, and the CCB Sign-Off Mark on every delivered dataset.

Protocol: The CCB Sign-Off Mark

How Ariana Nexus produces the data: the ADF Pipeline

Five pipeline stages, three capabilities, five validation gates. The ADF Pipeline produces the data; the Five-Gate Validation Protocol governs its quality, culture, provenance, and the conditions under which it was made.

The five stages

The Five Gates

01

Linguistic Accuracy — data linguistically accurate across 24 languages and dialects; inter-annotator agreement and gold-standard validation; the Pashto-Dari Parity Index applied.

02

Cultural Validity — data culturally valid, not merely literally correct; dialect-aware; cleared by the CCB Sign-Off Mark.

03

Standards Conformance — data-quality and annotation standards (ISO/IEC 5259, ISO/IEC 25012); RLHF and preference-data best practice; data provenance, consent, and licensing, with EU AI Act training-data transparency where applicable.

04

Population Risk — ethical labor: fair, living compensation, dignity, and never extractive; worker-wellbeing protection for sensitive or distressing content; consent and privacy; no rights-infringing data sourcing.

05

Institutional Sign-Off — datasets documented with provenance, lineage, quality metrics, and data cards — traceable and audit-ready.

The Four-Phase Delivery Cycle

Phase I

Situation — Understand

The model, the task (SFT, RLHF, evaluation), the target languages and dialects, and the data requirements mapped.

Cultural mapping · stakeholder calibration · constraint discovery

Phase II

Complication — Architect

The ADF Pipeline configured; annotation guidelines, dialect-aware reference design, and the quality framework built; the annotator cohort assembled ethically.

Program scaffolding · compliance baseline · governance charter

Phase III

Resolution — Deploy

Data produced — annotated, human feedback collected, dialect-aware reference sets built — with QA and provenance throughout.

In-context execution · data infrastructure

Phase IV

Measured Outcome — Govern

Data quality measured (agreement, gold-standard, downstream evaluation); data cards and provenance delivered; pipelines maintained and refreshed.

Continuous documentation · red-team validation · scheduled refresh

What governs AI training data — and where it stands.

The binding and voluntary regimes that reach training data, RLHF, and data labor in a healthcare context. Status as of June 2026.

InstrumentBodyStatusDate
HIPAA de-identification (Safe Harbor / Expert Determination)HHS OCRIn forceGuidance 2012
Section 1557 nondiscrimination — patient-care decision-support toolsHHS OCRIn forceApplied 1 May 2025
ONC HTI-1 — Predictive DSI transparency (FAVES, source attributes)ASTP / ONCIn forceCompliance 31 Dec 2024
Good Machine Learning Practice (GMLP) — representative datasets, data managementFDA · HC · MHRAVoluntary2021
Predetermined Change Control Plan (PCCP) for AI-enabled device softwareFDAFinal guidanceDec 2024
NIST AI RMF 1.0 — govern, map, measure, manage; data provenanceNISTVoluntary2023
EU AI Act, Art. 10 — data & data governance for high-risk AIEUApplies 2 Aug 2026Reg. 2024/1689
ISO/IEC 5259 — data quality for analytics & ML (Parts 1–4)ISO/IECPublished2024
ISO/IEC 8183 / 42001 / 23894 — data life cycle, AI management, AI riskISO/IECPublished2023

Voluntary frameworks (GMLP, PCCP principles, NIST AI RMF, ISO/IEC) are consensus standards, not law; HIPAA, Section 1557, ONC HTI-1, and the EU AI Act bind on their stated timelines. EU AI Act high-risk data-governance duties apply 2 August 2026 (the Act entered into force 1 August 2024); verify status at publish.

Healthcare

Where the data meets the clinic.

For health systems and their compliance officers, the training-data layer is now a regulated surface — a matter of civil rights, certification, and patient safety.

Section 1557 · in force 1 May 2025

Decision-support tools are in scope of civil-rights law.

The 2024 Section 1557 rule requires covered entities to identify and mitigate discrimination risk in patient-care decision-support tools — a duty that runs upstream into the data those tools learn from.

ONC HTI-1

Certification now demands training-data transparency.

HTI-1 requires certified health IT to surface 31 source attributes for predictive tools, so clinicians can judge whether a model is Fair, Appropriate, Valid, Effective, and Safe (FAVES) — including how it was trained.

Obermeyer, Science 2019

Bias is set in the label, not just the model.

The canonical clinical-AI failure traces to a cost-as-proxy training target, not architecture — proof that representative, well-specified data is itself a patient-safety control, not a downstream nicety.

HIPAA §164.514

PHI in training corpora is a hard prerequisite.

De-identification — Safe Harbor's 18 identifiers or documented Expert Determination — governs any PHI flowing into training data; unstructured clinical notes are the weak point, retaining residual identifiers.

What partnership looks like

1/4

Foundations

Scoped and mapped. The model, the task, the Afghan languages and dialects in scope, and the volume of data the work will need.

2/4

Activation

Built to standard. The pipeline configured, annotation guidelines written, and the Afghan annotator cohort assembled and paid fairly.

3/4

Operating Rhythm

The active state. Data produced with QA and provenance; quality measured against agreement and gold standards.

4/4

Continuous Stewardship

Refreshed and scaled. Pipelines maintained; data cards and provenance kept current as the model and task evolve.

Questions AI teams ask before they buy

What is Afghan language training data?

Afghan language training data is the labeled text, audio, and human-preference data used to train, fine-tune, and evaluate models in Afghan languages. Ariana Nexus produces Pashto training data, Dari training data, and sets in the other 22 Afghan languages — annotation, RLHF and preference data, instruction tuning, and dialect-aware evaluation references, each delivered with a data card.

Can a Farsi or Persian dataset be used for Dari?

Not safely. Afghanistan Dari and Iranian Persian differ in vocabulary, register, and idiom, and a model trained on Persian data produces Dari that Afghan users read as foreign or simply wrong. Ariana Nexus builds Afghanistan Dari data with Afghan native speakers, and keeps Kabuli, Herati, Badakhshani, and Hazaragi — a dialect of Dari — labeled separately rather than merged.

How much does Afghan language training data cost?

Cost depends on the language, the dialect coverage, the task, and the volume. An evaluation reference set of a few thousand items is a different order of work from a full RLHF preference program in Pashto and Dari. Ariana Nexus scopes each program against the model and the task before quoting, and publishes no rate card, because a rate quoted without scope is a guess.

Is Hazaragi a separate language or a Dari dialect?

Hazaragi is a dialect of Dari, not a separate language, although many buyers search for it by name. It carries distinct vocabulary and pronunciation, and an Afghan language dataset that merges it into standard Dari loses exactly the variation that matters. Ariana Nexus labels Hazaragi as a Dari variety and holds it separate in the data.

Should annotators be described as Pashto or Pashtun speakers?

Pashto is the language; Pashtun is the people who speak it, so the correct phrase is a Pashto annotator or a Pashto speaker. The same distinction applies to Afghan and Afghani: Afghan describes the people and the languages, while the afghani is Afghanistan's currency. Ariana Nexus uses the language names in every data card and annotation guideline.

How are Afghan annotators recruited and paid?

Ariana Nexus recruits Afghan native speakers from a graduate-level bench, matched to the dialect the work requires, and pays them fairly with worker-wellbeing protection for sensitive or distressing content. Consent, licensing, and annotation guidelines are documented for every dataset, so the provenance a regulator asks for is already on file.

Request an AI Data Factory Consultation.

Request a confidential briefing

Engagements are governed by the Ariana Nexus Trust Center.