LLM Evaluation & Red-Teaming in Afghan Languages
Ariana Nexus runs independent Afghan language LLM evaluation and red-teaming for AI labs and enterprise teams — accuracy, hallucination, jailbreak, refusal, and bias testing across 24 Afghan languages, scored by native speakers and documented to the NIST AI Risk Management Framework. Your model answers in fluent Pashto. That proves it is confident — not that it is correct.
Why does a model fail silently in Pashto and Dari?
A large language model answers a Pashto prompt with the same confidence whether it is right or catastrophically wrong — and the English benchmarks it passed say nothing about either. The failure is structural: the least training data, the fewest benchmarks, and the weakest safety filters sit in exactly the languages where confidence outruns competence.
The NIST AI Risk Management Framework asks every deployer to measure and manage that risk; the EU AI Act attaches conformity obligations to high-risk systems. None of it is satisfiable by an English benchmark for a model operating in 24 languages. Ariana Nexus validates and red-teams AI systems in the languages they actually serve — fluency separated from accuracy, guardrails tested where they are weakest, every finding documented for governance and disclosed for repair.
Fluency is not accuracy.
What the evidence shows — and your benchmark cannot
Every figure below is published and independently sourced. None of it is visible to an English-language evaluation.
Figures are published, peer-reviewed where applicable, and current as of June 2026. Full citations available on request.
The three teams behind an Afghan language LLM evaluation
Three institutional capabilities, working as one body of assurance you can put in front of an auditor.
Lived-expertise practitioners across all 24 Afghan languages — the cultural gatekeepers who keep every engagement anchored in ground truth, never extractive.
Native-speaker subject-matter experts across all 24 languages who read model output critically, catch culturally coded errors and hallucinations, and serve as expert human evaluators and red-teamers — the ground truth no automated metric replaces.
Protocol — The Cultural Hallucination AuditGoverned Afghan-language data infrastructure, evaluation benchmarks, and institutional-grade training assets meeting auditable standards.
Afghan-language evaluation datasets and benchmarks, the Cultural Hallucination Audit pipeline, multilingual evaluation harnesses, adversarial and red-team prompt sets, and human-in-the-loop scoring.
Protocol — The ADF PipelineAn audit-grade review regime translating cultural intelligence into compliance-ready practice — the governance layer threading through every engagement.
Methodology rigor and independence audit; reproducibility; bias and fairness review; responsible-disclosure governance for red-team findings; the CCB Sign-Off Mark on every assurance report.
Protocol — The CCB Sign-Off MarkHow Ariana Nexus exposes fluent-but-wrong output: the Cultural Hallucination Audit
The Cultural Hallucination Audit™ exposes fluent-but-wrong output; the Five-Gate Validation Protocol™ governs the assurance engagement, evaluation to red-team to sign-off.
Red-teaming is authorized and responsibly disclosed — findings go to the model's owner to remediate, never to weaponize, and we publish no exploit tooling.
Evaluation performed by people who can actually read the output — the ground truth no automated metric replaces.
Standards & compliance
Methodology, controls, and assurance documentation are open for review in the Ariana Nexus Trust Center →
What you receive.
Where your AI assurance actually stands
Five levels of validation maturity for multilingual AI. Most teams overestimate where they are.
Ariana Nexus engagements move a model from L0–L2 to L3 and hold it at L4.
Who leads the AI & Data Systems Practice

Model evaluation, multilingual AI systems, and assurance engineering · Cornell · University of Chicago.

Cultural compliance and responsible disclosure · Cornell University.
Published research & benchmarks

The Afghan Language AI Accuracy Report Card
How frontier and enterprise models actually perform across Afghan languages — graded against reference sets validated by the Cultural Compliance Bureau.
The Cultural Hallucination Audit™
Methodology for detecting fluent-but-wrong output.
The Low-Resource Safety Benchmark
Safety-guardrail robustness and jailbreak resistance across Afghan languages.
The Pashto–Dari Parity Index
Paired-language accuracy.
Questions AI teams ask before commissioning an evaluation
Ariana Nexus is a Washington, D.C.–area firm providing Afghan language services and cultural intelligence — interpretation, translation, cultural training, compliance support, and AI data — across 24 Afghan languages.
Why do LLMs fail more in Pashto and Dari than in English?
Because the training data is thin. Pashto and Dari sit among the lowest-resource languages a frontier model covers, so the model learns fluent surface form without the grounding to be correct — and it signals no less confidence when it is wrong. English benchmarks measure none of this.
Can automated metrics evaluate Afghan-language output?
Not on their own. BLEU, perplexity, and LLM-as-judge scoring inherit the same blind spot as the model, because they were built and calibrated on high-resource languages. Ariana Nexus scores output with native-speaker evaluators who can actually read it — including hallucination testing in Pashto — and documents the method so an auditor can reproduce the result.
Is Dari the same as Farsi for model evaluation?
No, and treating them as one is a common evaluation error. Afghanistan Dari and Iranian Persian differ in vocabulary, register, and idiom, so a model scored on Iranian Persian data can post a passing number while failing Afghan users. Ariana Nexus keeps Pashto LLM evaluation and Dari LLM evaluation as distinct tracks and never averages them into one figure.
What is Afghan language red teaming?
Adversarial safety testing carried out in the language itself rather than in English translation — low-resource language jailbreak testing, guardrail probing, refusal testing, and harmful-output testing in Pashto, Dari, and the other Afghan languages. Findings go to the model’s owner to remediate, never to weaponize, and Ariana Nexus publishes no exploit tooling.
Who provides Afghan language LLM evaluation?
Ariana Nexus provides Afghan language LLM evaluation and red-teaming for AI labs, enterprise deployers, and public institutions — accuracy, hallucination, jailbreak, refusal, and bias testing across 24 Afghan languages, scored by native speakers. Results are documented as NIST AI RMF multilingual evaluation evidence your governance function and auditors can use.
How much does an Afghan-language evaluation cost?
It depends on scope. Cost tracks how many Afghan languages are in scope, how many prompts and task types are evaluated, whether red-teaming and re-testing after remediation are included, and the depth of documentation required. Ariana Nexus scopes and prices per engagement and publishes no rate card.