Skip to main content

LLM Evaluation and Benchmark Development in Pashto, Dari, and 22 More Afghan Languages

Ariana Nexus builds the test sets, human evaluation panels and red-team programs that tell frontier AI labs and federal programs how their models actually perform in Pashto, Dari and 22 more Afghan languages. Native speakers write every item. Every score ships with the agreement statistics behind it.

24

Afghan languages in the coverage matrix

100%

of items double-annotated, then adjudicated

0

items machine-translated from an English benchmark

01

Key facts

Service
LLM evaluation, benchmark and test-set development, human preference data, and multilingual red teaming
Languages
Pashto and Dari first, extending to all 24 Afghan languages by readiness tier
Built for
Frontier and applied AI labs, federal agencies and prime contractors, model and data platforms, research institutions
Evaluators
Native speakers holding university degrees, qualified in the specific variety — not the macrolanguage
Delivered
Versioned dataset, datasheet, scoring harness, agreement report, findings memo, held-out split under separate custody
First engagement
A four to six week diagnostic on one language pair and one task family
Delivery
In-house, from Washington, D.C. No subcontractors, no brokered crowd panels

What is LLM evaluation in a low-resource language?

LLM evaluation in a low-resource language means measuring a model on tasks written and judged in that language by native speakers, instead of on English test sets pushed through machine translation. Benchmark development is the construction of that test set: the items, the reference answers, the scoring rubric, and the inter-annotator agreement evidence that makes the resulting scores reproducible and defensible.

02

Why English-first evaluation under-reports risk in Pashto and Dari

A model can pass a translated benchmark and still fail the language. Translation-based test sets measure the translation pipeline as much as the model, and the published record now shows how far apart the two numbers sit.

79%

of harmful prompts translated into low-resource languages drew an engaged, actionable response from GPT-4 on AdvBench — far above the rate for high- and mid-resource languages.

Yong, Menghini and Bach, Brown University (arXiv:2310.02446)

59.8% → 75.8%

average jailbreak rate across four low-resource languages once native-speaker human red teams replaced automated translation pipelines. Poor machine translation was suppressing the measured rate.

Multilingual jailbreaking of LLMs using low-resource languages, 2026 (arXiv:2605.18239)

64.6%

best published LLM accuracy on Belebele Pashto, a 900-item four-option comprehension set. Smaller models in the same family fall to 27.7%, and encoder similarity methods sit at chance.

PashtoCorp, 2026 (arXiv:2603.16354)

0.0%

Pashto language labels returned by Whisper Large V3 on verified Pashto text-to-speech audio — a language-identification failure a word-error-rate score alone never surfaces.

PashtoTTS-Bench, 2026 (arXiv:2605.26978)

  • Translated test sets measure the translator

    When an English benchmark is machine-translated into Pashto, a low score can mean the model failed or the translation failed, and the report cannot tell you which. The 2026 red-team result above quantifies the cost: automated translation suppressed the measured jailbreak rate by roughly sixteen points.

  • Dialect and variety collapse

    Afghanistan Dari is scored as Iranian Persian. Hazaragi, a variety of Dari with its own lexicon, is absent from the test set entirely. Afghan Turkmeni is evaluated against the standard Turkmen of Turkmenistan. Each substitution moves the score away from the population the model will actually serve.

  • Script fidelity failures that pass a metric

    Pashto-specific letters — ټ ډ ړ ږ ځ څ ڼ ښ — are dropped or normalized toward Urdu or Persian forms. Output can clear a word-error-rate threshold and still be unreadable to a Pashto speaker, because the metric never looked at the orthography.

  • Harm taxonomies written somewhere else

    Adversarial content in Pashto and Dari draws on the Afghan war, displacement, local political vocabulary and community-specific coded speech. An English harm taxonomy does not contain those categories, so the red team never writes the prompt that would have found the failure.

03

What we build

Six evaluation products. Most engagements begin with one and extend into the others as the measurement question sharpens.

  • 01

    Benchmark and test-set construction

    Native-authored items, reference answers, scoring rubrics, held-out splits and contamination controls. Built for your task families — reading comprehension, reasoning, instruction following, summarization, question answering, classification — not adapted from an English set.

    • Item authoring in-language
    • Held-out split under separate custody
    • Per-item provenance record
  • 02

    Human evaluation panels and preference data

    Pairwise preference collection, rubric-scored Likert evaluation, and preference datasets for reinforcement learning from human feedback. We also calibrate your LLM-as-a-judge against human labels and report where the judge and the panel diverge.

    • Pairwise and rubric scoring
    • RLHF preference sets
    • Judge-model calibration
  • 03

    Multilingual red teaming and adversarial evaluation

    Native-speaker adversarial prompt development, multi-turn attack construction, and harm taxonomy localization for AI safety evaluation. Findings go to the model owner to remediate — never to weaponize, and never as public exploit tooling.

    • Multi-turn attack sets
    • Localized harm taxonomy
    • Responsible disclosure only
  • 04

    Machine translation and cross-lingual quality evaluation

    MQM error typology scored by trained annotators, adequacy and fluency panels, post-edit distance, and validation of reference-free automatic metrics against human judgment in the specific variety.

    • MQM error typology
    • Adequacy and fluency panels
    • Automatic-metric validation
  • 05

    Speech, ASR and text-to-speech evaluation

    Word and character error rate by variety, script fidelity audit, language identification checks, and naturalness panels — the four-part reporting the published Pashto speech work shows a single intelligibility number cannot replace.

    • WER and CER by variety
    • Script fidelity audit
    • Language ID verification
  • 06

    Domain and cultural evaluation

    Clinical, legal, humanitarian, civic and conflict-era historical content, evaluated for factual accuracy, cultural appropriateness and refusal behavior — including how a model responds to Afghan users raising the Afghan war, displacement and moral trauma.

    • Clinical and legal domains
    • Refusal-behavior review
    • Cultural appropriateness
04

How we deliver

A six-phase method. The point of the sequence is that a number arrives with the evidence that makes it hold up under review.

  1. 01

    Scope the decision, not the dataset

    We start from the decision the score has to support — a release gate, a language expansion, a system-card claim, a contract acceptance test — and work backward to sample size, task families and agreement thresholds.

  2. 02

    Design the taxonomy and rubric in-language

    Rubrics are written in Pashto or Dari first and back-checked into English, so the categories match how the language behaves rather than how English does.

  3. 03

    Author items with named native speakers

    Degree-holding native speakers of the specific variety write every item. Source, author, rights status and consent are recorded per item at the point of authoring.

  4. 04

    Double-annotate blind, then adjudicate

    Every item is annotated independently twice. A third senior reviewer resolves each disagreement, and the adjudication is logged rather than silently overwritten.

  5. 05

    Validate statistically

    Krippendorff's alpha or Cohen's kappa per task, confidence intervals on every reported figure, and a written sample-size justification. Below the threshold set at scoping, the task goes back to rubric design.

  6. 06

    Deliver, then re-test on your cadence

    Versioned dataset, datasheet, scoring harness and findings memo. The held-out split stays under separate custody so your next release can be measured against an uncontaminated set.

05

The measurement standard

Every control below is evidenced in the delivery file. We do not publish a number we cannot show the working for.

Control
The standard we hold
How it is evidenced
ControlEvaluator qualification
The standard we holdUniversity degree held; documented proficiency in the specific variety claimed, not the macrolanguage
How it is evidencedQualification record per evaluator ID
ControlDouble annotation
The standard we hold100% of items annotated independently twice, blind to the first pass
How it is evidencedPer-item evaluator IDs in the delivery file
ControlAgreement
The standard we holdKrippendorff's alpha reported per task against a threshold agreed at scoping
How it is evidencedReliability report shipped with the dataset
ControlAdjudication
The standard we holdA third senior reviewer resolves every disagreement
How it is evidencedAdjudication log, decision by decision
ControlContamination control
The standard we holdItems authored, never scraped; held-out split never transmitted alongside the public split
How it is evidencedProvenance field on every item
ControlScript fidelity
The standard we holdPashto and Dari orthography checked against the variety standard, not normalized toward Urdu or Iranian Persian
How it is evidencedScript audit appended to the report
ControlProvenance and licensing
The standard we holdSource, author, rights status and consent recorded per item
How it is evidencedDatasheet in the standard research format
ControlContributor security
The standard we holdIdentities pseudonymized in deliverables; nothing routed through channels controlled by the de facto authorities in Afghanistan
How it is evidencedChain-of-custody record
06

Language coverage and readiness

Twenty-four Afghan languages across five families. We publish readiness tiers rather than a flat coverage claim, because speaker availability genuinely differs by two orders of magnitude across this list — and a vendor who tells you otherwise has not counted.

Tier
Languages
What we run today
TierTier 1
LanguagesPashto, Dari (including the Hazaragi variety)
What we run todayFull benchmark construction, preference panels, multi-turn red teaming, speech evaluation, re-test cadence
TierTier 2
LanguagesUzbeki, Turkmeni, Balochi, Pashayi, and the Nuristani group (Nuristani, Kati, Prasun, Waigali)
What we run todayBenchmark construction and human evaluation on an agreed scope, with panel size stated up front
TierTier 3
LanguagesAimaq, Ormuri, Parachi, Tirahi, Gawarbati, Kyrgyz, Brahui, and the Pamir group (Wakhi, Shughni, Sanglechi, Ishkashimi, Munji, Yidgha)
What we run todayTargeted evaluation sets and expert review, scoped to verified speaker availability rather than promised as volume
  • AimaqایماقTier 3
  • BalochiبلوچیTier 2
  • BrahuiبراهویTier 3
  • DariدریTier 1
  • GawarbatiگواربتیTier 3
  • HazaragiهزارگیTier 1
  • IshkashimiاشکاشمیTier 3
  • KatiکتیTier 2
  • KyrgyzقرغیزیTier 3
  • MunjiمنجیTier 3
  • NuristaniنورستانیTier 2
  • OrmuriارمړيTier 3
  • ParachiپراچیTier 3
  • Pashayiپشه‌ییTier 2
  • PashtoپښتوTier 1
  • PrasunپارونTier 2
  • SanglechiسنگلیچیTier 3
  • ShughniشغنیTier 3
  • TirahiتیراهیTier 3
  • TurkmeniترکمنیTier 2
  • UzbekiازبیکیTier 2
  • WaigaliویگلیTier 2
  • WakhiوخیTier 3
  • YidghaیدغهTier 3

Tier 1   Tier 2   Tier 3

Hazaragi is a variety of Dari and is listed as one; we staff it separately because the lexical distance is large enough to change a score. Dari deliverables are Afghanistan Dari, not Iranian Persian. Turkmeni deliverables are Afghan Turkmeni, not the standard Turkmen of Turkmenistan.

07

Governance, provenance and the regulatory record

Evaluation evidence is increasingly something a regulator or a contracting officer reads. We build it so it survives that reading.

Instrument
What it asks for
Status
InstrumentEU AI Act, Article 55 (GPAI with systemic risk)
What it asks forModel evaluation and adversarial testing, systemic-risk assessment and mitigation across the lifecycle
StatusCommission and AI Office enforcement powers in force since 2 August 2026
InstrumentEU AI Act, Article 53 with Annexes XI and XII
What it asks forTechnical documentation and downstream-provider information for general-purpose AI models
StatusApplied since 2 August 2025; models placed before that date must comply by 2 August 2027
InstrumentEU AI Act, Article 101
What it asks forPenalties for GPAI providers of up to €15 million or 3% of worldwide annual turnover
StatusAvailable to the Commission since 2 August 2026
InstrumentGPAI Code of Practice, Safety and Security chapter
What it asks forEvaluation, red teaming and reporting practice for signatory providers
StatusMore than 20 providers signed as of August 2026
InstrumentNIST AI Risk Management Framework and the Generative AI Profile
What it asks forMEASURE function: test, evaluation, verification and validation, including for non-English deployment
StatusVoluntary; routinely referenced in federal solicitations
InstrumentISO/IEC 42001 and ISO/IEC 23894
What it asks forAI management system and AI risk management practice
StatusMapped in delivery documentation on request
InstrumentOWASP Top 10 for LLM Applications and MITRE ATLAS
What it asks forAdversarial technique coverage and threat-informed test design
StatusUsed as the coverage checklist for red-team scope

Ariana Nexus produces evaluation evidence alongside your counsel, certifier and internal safety team — never in place of them. We do not certify conformity and we do not present a finding as a legal conclusion.

08

Who this is built for

  • Frontier and applied AI labs

    Pre-release evaluation, system-card and model-card evidence, and the data to decide whether a language is ready to claim.

  • Federal agencies and prime contractors

    Acceptance testing of language AI before it reaches mission use, with documentation written for a contracting file.

  • Model and data platform companies

    An Afghan-language evaluation bench you do not have to recruit, train and govern yourself.

  • Research institutions and benchmark consortia

    Co-authored, citable datasets with datasheets, licensing and reproducible scoring.

  • Health systems and humanitarian organizations

    Evaluation of AI already deployed to Afghan patients, Afghan refugees and Afghan diaspora communities before it carries clinical or protection consequences.

09

The people who write the items and sign the numbers

Ariana Nexus is not a crowd platform with a language filter on it. The people who design these evaluations hold degrees from Cornell, the University of Chicago, the University of British Columbia, the American University of Afghanistan and Otto-von-Guericke University — and they are from the communities whose languages they are measuring. That combination is the whole product: a linguist who has never read a reliability statistic and a statistician who has never spoken Pashto will each miss a different half of the failure.

Black-and-white architectural detail of a white faceted facade, used as a section image
Program oversight
  • Hassan Ukasha

    Managing Partner

    • B.S. Cornell University
    • M.P.H. Cornell University

    Oversees the firm's operations and this program: scope, evaluator qualification, custody of held-out material, and the evidence standard every published score is held to.

The delivery team

Zeba Haqbani

Senior Partner

  • B.A.
    The American University of Afghanistan
  • B.A.
    University of British Columbia

Builds the firm's institutional systems, technology and AI platforms; owns the evaluation harness and the delivery pipeline.

Hussain Ahmad

Principal

  • M.Eng.
    Cornell University
  • Ph.D.
    University of Chicago

AI and data engineering: rubric statistics, agreement modeling and judge-model calibration.

Wasil Peroz

Principal

  • B.A.
    Milli University
  • M.Sc.
    Otto-von-Guericke University Magdeburg

Institutional law; maps evaluation evidence onto the EU AI Act, the NIST framework and contract requirements.

Maryam Safi

Principal

  • B.A.
    Cornell University

Public-sector engagements: acceptance testing, federal delivery and documentation.

Behind the five names is the evaluation bench itself: native speakers of each variety, every one of them a university graduate, qualified against the variety they are claimed for and re-qualified on a schedule. Contributors are pseudonymized in every deliverable — an evaluator writing adversarial content in Pashto should never be identifiable from a dataset.

10

Why Ariana Nexus, and where we are different

  • 01

    We do not machine-translate a benchmark and call it coverage

    Every item in every deliverable is authored in the target language by a native speaker of the specific variety. This is the single largest source of error in published low-resource evaluation, and it is the one we removed first.

  • 02

    We resolve the variety, not the language code

    Afghanistan Dari is not Iranian Persian. Afghan Turkmeni is not the Turkmen of Turkmenistan. Hazaragi is a variety of Dari with a lexicon of its own. We staff, score and report at that resolution, because a model shipped to Afghan users will be judged at that resolution.

  • 03

    Scholars, not bilinguals

    We do not send you a bilingual. Evaluators hold university degrees; the people designing the instruments hold graduate degrees from leading universities. No U.S. certification exists for Afghan-language professionals, so Ariana Nexus wrote its own qualification standard, trains Afghan-language interpreters inside and outside the firm to it, and holds every evaluator on this bench to the same standard.

  • 04

    Every number arrives with its agreement statistics

    Alpha per task, confidence intervals, sample-size justification, adjudication log. A score without its reliability evidence is an opinion, and an opinion is not what you are buying.

  • 05

    One firm, one point of accountability

    Every part of the engagement is produced by our own people. No subcontractors, no brokered specialists, no crowd marketplace behind the curtain — which also means no unexplained drift in who is labeling your data.

  • 06

    A compliance line we do not cross

    Ariana Nexus routes no documents, data or inquiries through channels controlled by the de facto authorities in Afghanistan. Contributors are diaspora-based and pseudonymized. For adversarial work, that is a safety requirement, not a preference.

  • 07

    We publish, and we are citable

    The firm's language research is published under its own name so that a reviewer, a regulator or an AI search system can trace a claim back to a document rather than to a sales page.

Ariana Nexus
General annotation platform
Machine-translated test set
Who writes the items
Ariana NexusDegree-holding native speakers of the named variety
General annotation platformWhoever passes a generic language screen
Machine-translated test setA translation system
Variety resolution
Ariana NexusAfghanistan Dari, Afghan Turkmeni, Hazaragi staffed separately
General annotation platformMacrolanguage code only
Machine-translated test setMacrolanguage code only
Script fidelity check
Ariana NexusAudited against the variety standard
General annotation platformNot performed
Machine-translated test setNot performed
Agreement reporting
Ariana NexusAlpha per task, intervals, adjudication log
General annotation platformOccasionally, on request
Machine-translated test setNot applicable
Adversarial coverage
Ariana NexusLocalized harm taxonomy, multi-turn, native-authored
General annotation platformEnglish taxonomy, translated
Machine-translated test setEnglish taxonomy, translated
Provenance and licensing
Ariana NexusPer-item source, rights and consent
General annotation platformPlatform-level terms
Machine-translated test setInherited from the source set
Accountability
Ariana NexusOne firm, named partners, in-house delivery
General annotation platformA marketplace of contributors
Machine-translated test setNone
11

What you receive and engagement models

  • Versioned datasetPer-item provenance, evaluator IDs, license terms and variety tags.
  • Datasheet and READMEComposition, collection process, intended use, limitations and known biases, in the standard research format.
  • Scoring harnessScripts that run against your evaluation stack, with reference outputs for verification.
  • Reliability reportAgreement per task, confidence intervals, sample-size justification, adjudication log.
  • Findings memoThe failure taxonomy the evaluation exposed, ranked by severity and frequency, with worked examples.
  • Held-out splitRetained under separate custody so your next checkpoint can be measured against an uncontaminated set.
  • Re-test scheduleThe same instrument re-run on your release cadence, so the numbers are comparable over time.

Engagement models

  • Diagnostic

    Four to six weeks

    One language pair, one task family, one report. For teams deciding whether the gap they suspect is real and how large it is.

  • Benchmark program

    Three to six months

    Full test-set construction across task families, with adversarial coverage, a reliability standard agreed at scoping and a re-test schedule.

  • Standing evaluation panel

    Annual

    A named panel on retainer, scoped to your release cadence, with defined turnaround for incident response and pre-release gates.

Ariana Nexus does not compete on price per label. Engagements carry a minimum scope, because the evidence standard above does not scale down — an agreement statistic computed on too few items is worse than no statistic at all.

12

Questions buyers ask

What is LLM evaluation in a low-resource language like Pashto or Dari?

LLM evaluation in a low-resource language means measuring a model on tasks written and judged in that language by native speakers, instead of on English test sets pushed through machine translation. Benchmark development is the construction of that test set: the items, the reference answers, the scoring rubric, and the inter-annotator agreement evidence that makes the resulting scores reproducible and defensible.

Why can't we just translate our English benchmark into Pashto?

Because the resulting score measures the translation system as much as the model, and you cannot separate the two afterward. Published 2026 work found that replacing automated translation with native-speaker human red teams raised the measured jailbreak rate from 59.8% to 75.8% across four low-resource languages — the translated pipeline was hiding roughly sixteen points of real vulnerability.

Is Dari the same as Persian or Farsi?

No. Dari is the Afghan variety of Persian, and it differs from Iranian Persian in vocabulary, register and a great deal of institutional and clinical terminology. A model evaluated on Iranian Persian data has not been evaluated for Afghan users. All Ariana Nexus Dari work is Afghanistan Dari.

Do you treat Hazaragi as a separate language?

Hazaragi is a variety of Dari and we describe it that way. We staff it separately, because its lexical distance from standard Dari is large enough to change a score, and Hazara speakers are a population your deployment will meet.

How many Afghan languages can you actually evaluate?

All 24 appear in our coverage matrix, in three published readiness tiers. Pashto and Dari run at full benchmark, red-team and speech capability today. Uzbeki, Turkmeni, Balochi, Pashayi and the Nuristani group run benchmark construction and human evaluation on an agreed scope. The smallest languages — Ormuri, Parachi, Tirahi and the Pamir group — get targeted sets scoped to verified speaker availability. We publish the tiers rather than promise uniform volume across all 24.

Who writes the test items?

Native speakers of the specific variety who hold university degrees, working to a rubric written in the target language. Each item carries the author's evaluator ID, its source, rights status and consent record.

How do you report reliability?

Krippendorff's alpha or Cohen's kappa per task, with confidence intervals and a written sample-size justification. Every item is double-annotated blind and every disagreement is adjudicated by a third senior reviewer, with the decision logged.

Do you run multilingual red teaming?

Yes — native-speaker adversarial prompt development, multi-turn attack construction and harm taxonomy localization for Pashto and Dari, with Tier 2 languages on scope. Findings go to the model owner to remediate. We do not publish exploit tooling and we do not release attack sets publicly.

Can your evidence support EU AI Act or NIST documentation?

It is built for it. Deliverables map onto Article 55 model-evaluation and adversarial-testing duties, Article 53 technical documentation, and the MEASURE function of the NIST AI Risk Management Framework and its Generative AI Profile. We produce the evidence alongside your counsel and certifier, never in place of them.

How do you handle data provenance, licensing and contributor safety?

Source, author, rights status and consent are recorded per item, and the dataset ships with a datasheet. Contributors are pseudonymized in every deliverable. Ariana Nexus routes no data or inquiries through channels controlled by the de facto authorities in Afghanistan.

Can you evaluate speech models — ASR and text-to-speech — in Pashto?

Yes. Word and character error rate by variety, plus the three checks a single intelligibility number misses: script fidelity, language identification and native naturalness judgment. Published 2026 work found a leading ASR system returning zero Pashto language labels on verified Pashto audio, which is exactly the class of failure a WER-only report hides.

How long does a first engagement take, and what does it cost?

A diagnostic runs four to six weeks. Pricing is scoped to the task families, the number of items and the agreement threshold you need, and engagements carry a minimum scope. Ariana Nexus is not a low-cost labeling option and does not price per label.

Brief a partner on the model you need measured

Tell us the model, the languages and the decision the score has to support. You will speak with a partner, under NDA, not a sales desk.

Start a conversation

Ariana Nexus · 1717 Pennsylvania Avenue NW, 10th Floor, Washington, D.C. 20006

Related services

Reviewed by Hassan Ukasha, Managing Partner · Last reviewed 20 September 2026 · AN·SVC·2026·TECH·01