LLM Evaluation and Benchmark Development in Pashto, Dari, and 22 More Afghan Languages
Ariana Nexus builds the test sets, human evaluation panels and red-team programs that tell frontier AI labs and federal programs how their models actually perform in Pashto, Dari and 22 more Afghan languages. Native speakers write every item. Every score ships with the agreement statistics behind it.
Afghan languages in the coverage matrix
of items double-annotated, then adjudicated
items machine-translated from an English benchmark
Key facts
- Service
- LLM evaluation, benchmark and test-set development, human preference data, and multilingual red teaming
- Languages
- Pashto and Dari first, extending to all 24 Afghan languages by readiness tier
- Built for
- Frontier and applied AI labs, federal agencies and prime contractors, model and data platforms, research institutions
- Evaluators
- Native speakers holding university degrees, qualified in the specific variety — not the macrolanguage
- Delivered
- Versioned dataset, datasheet, scoring harness, agreement report, findings memo, held-out split under separate custody
- First engagement
- A four to six week diagnostic on one language pair and one task family
- Delivery
- In-house, from Washington, D.C. No subcontractors, no brokered crowd panels
What is LLM evaluation in a low-resource language?
LLM evaluation in a low-resource language means measuring a model on tasks written and judged in that language by native speakers, instead of on English test sets pushed through machine translation. Benchmark development is the construction of that test set: the items, the reference answers, the scoring rubric, and the inter-annotator agreement evidence that makes the resulting scores reproducible and defensible.
Why English-first evaluation under-reports risk in Pashto and Dari
A model can pass a translated benchmark and still fail the language. Translation-based test sets measure the translation pipeline as much as the model, and the published record now shows how far apart the two numbers sit.
of harmful prompts translated into low-resource languages drew an engaged, actionable response from GPT-4 on AdvBench — far above the rate for high- and mid-resource languages.
Yong, Menghini and Bach, Brown University (arXiv:2310.02446)
average jailbreak rate across four low-resource languages once native-speaker human red teams replaced automated translation pipelines. Poor machine translation was suppressing the measured rate.
Multilingual jailbreaking of LLMs using low-resource languages, 2026 (arXiv:2605.18239)
best published LLM accuracy on Belebele Pashto, a 900-item four-option comprehension set. Smaller models in the same family fall to 27.7%, and encoder similarity methods sit at chance.
PashtoCorp, 2026 (arXiv:2603.16354)
Pashto language labels returned by Whisper Large V3 on verified Pashto text-to-speech audio — a language-identification failure a word-error-rate score alone never surfaces.
PashtoTTS-Bench, 2026 (arXiv:2605.26978)
Translated test sets measure the translator
When an English benchmark is machine-translated into Pashto, a low score can mean the model failed or the translation failed, and the report cannot tell you which. The 2026 red-team result above quantifies the cost: automated translation suppressed the measured jailbreak rate by roughly sixteen points.
Dialect and variety collapse
Afghanistan Dari is scored as Iranian Persian. Hazaragi, a variety of Dari with its own lexicon, is absent from the test set entirely. Afghan Turkmeni is evaluated against the standard Turkmen of Turkmenistan. Each substitution moves the score away from the population the model will actually serve.
Script fidelity failures that pass a metric
Pashto-specific letters — ټ ډ ړ ږ ځ څ ڼ ښ — are dropped or normalized toward Urdu or Persian forms. Output can clear a word-error-rate threshold and still be unreadable to a Pashto speaker, because the metric never looked at the orthography.
Harm taxonomies written somewhere else
Adversarial content in Pashto and Dari draws on the Afghan war, displacement, local political vocabulary and community-specific coded speech. An English harm taxonomy does not contain those categories, so the red team never writes the prompt that would have found the failure.
What we build
Six evaluation products. Most engagements begin with one and extend into the others as the measurement question sharpens.
- 01
Benchmark and test-set construction
Native-authored items, reference answers, scoring rubrics, held-out splits and contamination controls. Built for your task families — reading comprehension, reasoning, instruction following, summarization, question answering, classification — not adapted from an English set.
- 02
Human evaluation panels and preference data
Pairwise preference collection, rubric-scored Likert evaluation, and preference datasets for reinforcement learning from human feedback. We also calibrate your LLM-as-a-judge against human labels and report where the judge and the panel diverge.
- 03
Multilingual red teaming and adversarial evaluation
Native-speaker adversarial prompt development, multi-turn attack construction, and harm taxonomy localization for AI safety evaluation. Findings go to the model owner to remediate — never to weaponize, and never as public exploit tooling.
- 04
Machine translation and cross-lingual quality evaluation
MQM error typology scored by trained annotators, adequacy and fluency panels, post-edit distance, and validation of reference-free automatic metrics against human judgment in the specific variety.
- 05
Speech, ASR and text-to-speech evaluation
Word and character error rate by variety, script fidelity audit, language identification checks, and naturalness panels — the four-part reporting the published Pashto speech work shows a single intelligibility number cannot replace.
- 06
Domain and cultural evaluation
Clinical, legal, humanitarian, civic and conflict-era historical content, evaluated for factual accuracy, cultural appropriateness and refusal behavior — including how a model responds to Afghan users raising the Afghan war, displacement and moral trauma.
How we deliver
A six-phase method. The point of the sequence is that a number arrives with the evidence that makes it hold up under review.
- 01
Scope the decision, not the dataset
We start from the decision the score has to support — a release gate, a language expansion, a system-card claim, a contract acceptance test — and work backward to sample size, task families and agreement thresholds.
- 02
Design the taxonomy and rubric in-language
Rubrics are written in Pashto or Dari first and back-checked into English, so the categories match how the language behaves rather than how English does.
- 03
Author items with named native speakers
Degree-holding native speakers of the specific variety write every item. Source, author, rights status and consent are recorded per item at the point of authoring.
- 04
Double-annotate blind, then adjudicate
Every item is annotated independently twice. A third senior reviewer resolves each disagreement, and the adjudication is logged rather than silently overwritten.
- 05
Validate statistically
Krippendorff's alpha or Cohen's kappa per task, confidence intervals on every reported figure, and a written sample-size justification. Below the threshold set at scoping, the task goes back to rubric design.
- 06
Deliver, then re-test on your cadence
Versioned dataset, datasheet, scoring harness and findings memo. The held-out split stays under separate custody so your next release can be measured against an uncontaminated set.
The measurement standard
Every control below is evidenced in the delivery file. We do not publish a number we cannot show the working for.
Language coverage and readiness
Twenty-four Afghan languages across five families. We publish readiness tiers rather than a flat coverage claim, because speaker availability genuinely differs by two orders of magnitude across this list — and a vendor who tells you otherwise has not counted.
- AimaqایماقTier 3
- BalochiبلوچیTier 2
- BrahuiبراهویTier 3
- DariدریTier 1
- GawarbatiگواربتیTier 3
- HazaragiهزارگیTier 1
- IshkashimiاشکاشمیTier 3
- KatiکتیTier 2
- KyrgyzقرغیزیTier 3
- MunjiمنجیTier 3
- NuristaniنورستانیTier 2
- OrmuriارمړيTier 3
- ParachiپراچیTier 3
- PashayiپشهییTier 2
- PashtoپښتوTier 1
- PrasunپارونTier 2
- SanglechiسنگلیچیTier 3
- ShughniشغنیTier 3
- TirahiتیراهیTier 3
- TurkmeniترکمنیTier 2
- UzbekiازبیکیTier 2
- WaigaliویگلیTier 2
- WakhiوخیTier 3
- YidghaیدغهTier 3
Tier 1 Tier 2 Tier 3
Hazaragi is a variety of Dari and is listed as one; we staff it separately because the lexical distance is large enough to change a score. Dari deliverables are Afghanistan Dari, not Iranian Persian. Turkmeni deliverables are Afghan Turkmeni, not the standard Turkmen of Turkmenistan.
Governance, provenance and the regulatory record
Evaluation evidence is increasingly something a regulator or a contracting officer reads. We build it so it survives that reading.
Ariana Nexus produces evaluation evidence alongside your counsel, certifier and internal safety team — never in place of them. We do not certify conformity and we do not present a finding as a legal conclusion.
Who this is built for
Frontier and applied AI labs
Pre-release evaluation, system-card and model-card evidence, and the data to decide whether a language is ready to claim.
Federal agencies and prime contractors
Acceptance testing of language AI before it reaches mission use, with documentation written for a contracting file.
Model and data platform companies
An Afghan-language evaluation bench you do not have to recruit, train and govern yourself.
Research institutions and benchmark consortia
Co-authored, citable datasets with datasheets, licensing and reproducible scoring.
Health systems and humanitarian organizations
Evaluation of AI already deployed to Afghan patients, Afghan refugees and Afghan diaspora communities before it carries clinical or protection consequences.
The people who write the items and sign the numbers
Ariana Nexus is not a crowd platform with a language filter on it. The people who design these evaluations hold degrees from Cornell, the University of Chicago, the University of British Columbia, the American University of Afghanistan and Otto-von-Guericke University — and they are from the communities whose languages they are measuring. That combination is the whole product: a linguist who has never read a reliability statistic and a statistician who has never spoken Pashto will each miss a different half of the failure.

Hassan Ukasha
Managing Partner
- B.S. Cornell University
- M.P.H. Cornell University
Oversees the firm's operations and this program: scope, evaluator qualification, custody of held-out material, and the evidence standard every published score is held to.
Zeba Haqbani
Senior Partner
- B.A.The American University of Afghanistan
- B.A.University of British Columbia
Builds the firm's institutional systems, technology and AI platforms; owns the evaluation harness and the delivery pipeline.
Hussain Ahmad
Principal
- M.Eng.Cornell University
- Ph.D.University of Chicago
AI and data engineering: rubric statistics, agreement modeling and judge-model calibration.
Wasil Peroz
Principal
- B.A.Milli University
- M.Sc.Otto-von-Guericke University Magdeburg
Institutional law; maps evaluation evidence onto the EU AI Act, the NIST framework and contract requirements.
Maryam Safi
Principal
- B.A.Cornell University
Public-sector engagements: acceptance testing, federal delivery and documentation.
Behind the five names is the evaluation bench itself: native speakers of each variety, every one of them a university graduate, qualified against the variety they are claimed for and re-qualified on a schedule. Contributors are pseudonymized in every deliverable — an evaluator writing adversarial content in Pashto should never be identifiable from a dataset.
Why Ariana Nexus, and where we are different
- 01
We do not machine-translate a benchmark and call it coverage
Every item in every deliverable is authored in the target language by a native speaker of the specific variety. This is the single largest source of error in published low-resource evaluation, and it is the one we removed first.
- 02
We resolve the variety, not the language code
Afghanistan Dari is not Iranian Persian. Afghan Turkmeni is not the Turkmen of Turkmenistan. Hazaragi is a variety of Dari with a lexicon of its own. We staff, score and report at that resolution, because a model shipped to Afghan users will be judged at that resolution.
- 03
Scholars, not bilinguals
We do not send you a bilingual. Evaluators hold university degrees; the people designing the instruments hold graduate degrees from leading universities. No U.S. certification exists for Afghan-language professionals, so Ariana Nexus wrote its own qualification standard, trains Afghan-language interpreters inside and outside the firm to it, and holds every evaluator on this bench to the same standard.
- 04
Every number arrives with its agreement statistics
Alpha per task, confidence intervals, sample-size justification, adjudication log. A score without its reliability evidence is an opinion, and an opinion is not what you are buying.
- 05
One firm, one point of accountability
Every part of the engagement is produced by our own people. No subcontractors, no brokered specialists, no crowd marketplace behind the curtain — which also means no unexplained drift in who is labeling your data.
- 06
A compliance line we do not cross
Ariana Nexus routes no documents, data or inquiries through channels controlled by the de facto authorities in Afghanistan. Contributors are diaspora-based and pseudonymized. For adversarial work, that is a safety requirement, not a preference.
- 07
We publish, and we are citable
The firm's language research is published under its own name so that a reviewer, a regulator or an AI search system can trace a claim back to a document rather than to a sales page.
What you receive and engagement models
- Versioned datasetPer-item provenance, evaluator IDs, license terms and variety tags.
- Datasheet and READMEComposition, collection process, intended use, limitations and known biases, in the standard research format.
- Scoring harnessScripts that run against your evaluation stack, with reference outputs for verification.
- Reliability reportAgreement per task, confidence intervals, sample-size justification, adjudication log.
- Findings memoThe failure taxonomy the evaluation exposed, ranked by severity and frequency, with worked examples.
- Held-out splitRetained under separate custody so your next checkpoint can be measured against an uncontaminated set.
- Re-test scheduleThe same instrument re-run on your release cadence, so the numbers are comparable over time.
Engagement models
Diagnostic
Four to six weeks
One language pair, one task family, one report. For teams deciding whether the gap they suspect is real and how large it is.
Benchmark program
Three to six months
Full test-set construction across task families, with adversarial coverage, a reliability standard agreed at scoping and a re-test schedule.
Standing evaluation panel
Annual
A named panel on retainer, scoped to your release cadence, with defined turnaround for incident response and pre-release gates.
Ariana Nexus does not compete on price per label. Engagements carry a minimum scope, because the evidence standard above does not scale down — an agreement statistic computed on too few items is worse than no statistic at all.
Questions buyers ask
What is LLM evaluation in a low-resource language like Pashto or Dari?
LLM evaluation in a low-resource language means measuring a model on tasks written and judged in that language by native speakers, instead of on English test sets pushed through machine translation. Benchmark development is the construction of that test set: the items, the reference answers, the scoring rubric, and the inter-annotator agreement evidence that makes the resulting scores reproducible and defensible.
Why can't we just translate our English benchmark into Pashto?
Because the resulting score measures the translation system as much as the model, and you cannot separate the two afterward. Published 2026 work found that replacing automated translation with native-speaker human red teams raised the measured jailbreak rate from 59.8% to 75.8% across four low-resource languages — the translated pipeline was hiding roughly sixteen points of real vulnerability.
Is Dari the same as Persian or Farsi?
No. Dari is the Afghan variety of Persian, and it differs from Iranian Persian in vocabulary, register and a great deal of institutional and clinical terminology. A model evaluated on Iranian Persian data has not been evaluated for Afghan users. All Ariana Nexus Dari work is Afghanistan Dari.
Do you treat Hazaragi as a separate language?
Hazaragi is a variety of Dari and we describe it that way. We staff it separately, because its lexical distance from standard Dari is large enough to change a score, and Hazara speakers are a population your deployment will meet.
How many Afghan languages can you actually evaluate?
All 24 appear in our coverage matrix, in three published readiness tiers. Pashto and Dari run at full benchmark, red-team and speech capability today. Uzbeki, Turkmeni, Balochi, Pashayi and the Nuristani group run benchmark construction and human evaluation on an agreed scope. The smallest languages — Ormuri, Parachi, Tirahi and the Pamir group — get targeted sets scoped to verified speaker availability. We publish the tiers rather than promise uniform volume across all 24.
Who writes the test items?
Native speakers of the specific variety who hold university degrees, working to a rubric written in the target language. Each item carries the author's evaluator ID, its source, rights status and consent record.
How do you report reliability?
Krippendorff's alpha or Cohen's kappa per task, with confidence intervals and a written sample-size justification. Every item is double-annotated blind and every disagreement is adjudicated by a third senior reviewer, with the decision logged.
Do you run multilingual red teaming?
Yes — native-speaker adversarial prompt development, multi-turn attack construction and harm taxonomy localization for Pashto and Dari, with Tier 2 languages on scope. Findings go to the model owner to remediate. We do not publish exploit tooling and we do not release attack sets publicly.
Can your evidence support EU AI Act or NIST documentation?
It is built for it. Deliverables map onto Article 55 model-evaluation and adversarial-testing duties, Article 53 technical documentation, and the MEASURE function of the NIST AI Risk Management Framework and its Generative AI Profile. We produce the evidence alongside your counsel and certifier, never in place of them.
How do you handle data provenance, licensing and contributor safety?
Source, author, rights status and consent are recorded per item, and the dataset ships with a datasheet. Contributors are pseudonymized in every deliverable. Ariana Nexus routes no data or inquiries through channels controlled by the de facto authorities in Afghanistan.
Can you evaluate speech models — ASR and text-to-speech — in Pashto?
Yes. Word and character error rate by variety, plus the three checks a single intelligibility number misses: script fidelity, language identification and native naturalness judgment. Published 2026 work found a leading ASR system returning zero Pashto language labels on verified Pashto audio, which is exactly the class of failure a WER-only report hides.
How long does a first engagement take, and what does it cost?
A diagnostic runs four to six weeks. Pricing is scoped to the task families, the number of items and the agreement threshold you need, and engagements carry a minimum scope. Ariana Nexus is not a low-cost labeling option and does not price per label.
Brief a partner on the model you need measured
Tell us the model, the languages and the decision the score has to support. You will speak with a partner, under NDA, not a sales desk.
Start a conversationAriana Nexus · 1717 Pennsylvania Avenue NW, 10th Floor, Washington, D.C. 20006
Related services
- AI and machine translation quality review
Human review of machine-translated Pashto and Dari output before it reaches the public.
- Afghan language data and annotation programs
Training, fine-tuning and preference corpora built under the same provenance standard.
- AI governance and compliance advisory
Mapping model evidence onto the EU AI Act, the NIST framework and contract requirements.
Reviewed by Hassan Ukasha, Managing Partner · Last reviewed 20 September 2026 · AN·SVC·2026·TECH·01