LLM Evaluation & Red-Teaming in Afghan Languages

Ariana Nexus runs independent Afghan language LLM evaluation and red-teaming for AI labs and enterprise teams — accuracy, hallucination, jailbreak, refusal, and bias testing across 24 Afghan languages, scored by native speakers and documented to the NIST AI Risk Management Framework. Your model answers in fluent Pashto. That proves it is confident — not that it is correct.

Why does a model fail silently in Pashto and Dari?

A large language model answers a Pashto prompt with the same confidence whether it is right or catastrophically wrong — and the English benchmarks it passed say nothing about either. The failure is structural: the least training data, the fewest benchmarks, and the weakest safety filters sit in exactly the languages where confidence outruns competence.

The NIST AI Risk Management Framework asks every deployer to measure and manage that risk; the EU AI Act attaches conformity obligations to high-risk systems. None of it is satisfiable by an English benchmark for a model operating in 24 languages. Ariana Nexus validates and red-teams AI systems in the languages they actually serve — fluency separated from accuracy, guardrails tested where they are weakest, every finding documented for governance and disclosed for repair.

FluencyAccuracyHigh-resourceLow-resourcedivergence
The gap no benchmark sees — fluent output, in a language the evaluation never tested.
24Afghan languages evaluated — most of them low-resource, where models fail and no one checks.
NIST AI RMFMeasure and Manage, in the languages your benchmarks skip.
Sovereign Speech IndexThe Ariana Nexus model-performance measure.

Fluency is not accuracy.

What the evidence shows — and your benchmark cannot

Every figure below is published and independently sourced. None of it is visible to an English-language evaluation.

FindingWhat it meansSource
79%Unsafe prompts translated into low-resource languages bypassed a frontier model's guardrails 79% of the time — versus under 1% in English. Safety training does not transfer across languages.Yong, Menghini and Bach, 2023–24
44%Even best-case English clinical summarization hallucinated — 44% of those hallucinations were clinically significant.Asgari et al., npj Digital Medicine, 2025
<45%For the most under-resourced languages — the category Pashto and Dari fall into — most frontier models score below 45% accuracy.IndicParam, 2026

Figures are published, peer-reviewed where applicable, and current as of June 2026. Full citations available on request.

Operating Model

The three teams behind an Afghan language LLM evaluation

Three institutional capabilities, working as one body of assurance you can put in front of an auditor.

HIC
Human Intelligence Collective

Lived-expertise practitioners across all 24 Afghan languages — the cultural gatekeepers who keep every engagement anchored in ground truth, never extractive.

Native-speaker subject-matter experts across all 24 languages who read model output critically, catch culturally coded errors and hallucinations, and serve as expert human evaluators and red-teamers — the ground truth no automated metric replaces.

Protocol — The Cultural Hallucination Audit
ADF
AI Data Factory

Governed Afghan-language data infrastructure, evaluation benchmarks, and institutional-grade training assets meeting auditable standards.

Afghan-language evaluation datasets and benchmarks, the Cultural Hallucination Audit pipeline, multilingual evaluation harnesses, adversarial and red-team prompt sets, and human-in-the-loop scoring.

Protocol — The ADF Pipeline
CCB
Cultural Compliance Bureau

An audit-grade review regime translating cultural intelligence into compliance-ready practice — the governance layer threading through every engagement.

Methodology rigor and independence audit; reproducibility; bias and fairness review; responsible-disclosure governance for red-team findings; the CCB Sign-Off Mark on every assurance report.

Protocol — The CCB Sign-Off Mark
HICADFCCBOne governed assurance report

How Ariana Nexus exposes fluent-but-wrong output: the Cultural Hallucination Audit

The Cultural Hallucination Audit™ exposes fluent-but-wrong output; the Five-Gate Validation Protocol™ governs the assurance engagement, evaluation to red-team to sign-off.

Red-teaming is authorized and responsibly disclosed — findings go to the model's owner to remediate, never to weaponize, and we publish no exploit tooling.

Native-speaker assurance

Evaluation performed by people who can actually read the output — the ground truth no automated metric replaces.

What Partnership Looks Like

What you receive.

model opacityauditable truth
Independent multilingual evaluation across 24 Afghan languages.Accuracy, fluency, and reliability scored by native-speaker experts — not automated metrics alone.
A Cultural Hallucination Audit.Fluent-but-wrong output and culturally coded errors detected, where standard evaluation sees nothing.
AI red-teaming for low-resource languages.Safety-guardrail robustness, jailbreak resistance, bias, and harmful-output testing — responsibly disclosed for remediation.
NIST AI RMF-aligned assurance documentation.Measure-and-Manage evidence your governance and auditors can use, reproducible, with documented independence.
A Sovereign Speech Index score and a Low-Resource Safety Benchmark result.Performance and guardrail robustness across the 24 languages, benchmarked.
Re-testing after remediation, and lifecycle monitoring.Assurance that continues as the model evolves — not a one-time certificate.

Where your AI assurance actually stands

Five levels of validation maturity for multilingual AI. Most teams overestimate where they are.

L0
Self-attested
Vendor claims and English benchmarks. No independent evidence in the languages served.
L1
English-validated
Tested in English; low-resource-language behavior is unknown and unmeasured.
L2
Spot-checked Most teams sit here
Ad hoc in-language review, not reproducible and not documented for governance.
L3
Independently validated
Native-speaker evaluation and red-teaming, documented to the NIST AI Risk Management Framework.
L4
Continuously assured
Re-tested after remediation, benchmarked, and monitored across the model lifecycle. Audit-ready.

Ariana Nexus engagements move a model from L0–L2 to L3 and hold it at L4.

Who Leads

Who leads the AI & Data Systems Practice

Hussain Ahmad, Senior Practice Leader, AI & Data Engineering, Ariana Nexus
Hussain Ahmad
Senior Practice Leader, AI & Data Engineering

Model evaluation, multilingual AI systems, and assurance engineering · Cornell · University of Chicago.

Maryam Safi, Principal, Healthcare Research, Ariana Nexus
Maryam Safi
Principal, Healthcare Research

Cultural compliance and responsible disclosure · Cornell University.

Questions AI teams ask before commissioning an evaluation

Ariana Nexus is a Washington, D.C.–area firm providing Afghan language services and cultural intelligence — interpretation, translation, cultural training, compliance support, and AI data — across 24 Afghan languages.

Why do LLMs fail more in Pashto and Dari than in English?

Because the training data is thin. Pashto and Dari sit among the lowest-resource languages a frontier model covers, so the model learns fluent surface form without the grounding to be correct — and it signals no less confidence when it is wrong. English benchmarks measure none of this.

Can automated metrics evaluate Afghan-language output?

Not on their own. BLEU, perplexity, and LLM-as-judge scoring inherit the same blind spot as the model, because they were built and calibrated on high-resource languages. Ariana Nexus scores output with native-speaker evaluators who can actually read it — including hallucination testing in Pashto — and documents the method so an auditor can reproduce the result.

Is Dari the same as Farsi for model evaluation?

No, and treating them as one is a common evaluation error. Afghanistan Dari and Iranian Persian differ in vocabulary, register, and idiom, so a model scored on Iranian Persian data can post a passing number while failing Afghan users. Ariana Nexus keeps Pashto LLM evaluation and Dari LLM evaluation as distinct tracks and never averages them into one figure.

What is Afghan language red teaming?

Adversarial safety testing carried out in the language itself rather than in English translation — low-resource language jailbreak testing, guardrail probing, refusal testing, and harmful-output testing in Pashto, Dari, and the other Afghan languages. Findings go to the model’s owner to remediate, never to weaponize, and Ariana Nexus publishes no exploit tooling.

Who provides Afghan language LLM evaluation?

Ariana Nexus provides Afghan language LLM evaluation and red-teaming for AI labs, enterprise deployers, and public institutions — accuracy, hallucination, jailbreak, refusal, and bias testing across 24 Afghan languages, scored by native speakers. Results are documented as NIST AI RMF multilingual evaluation evidence your governance function and auditors can use.

How much does an Afghan-language evaluation cost?

It depends on scope. Cost tracks how many Afghan languages are in scope, how many prompts and task types are evaluated, whether red-teaming and re-testing after remediation are included, and the depth of documentation required. Ariana Nexus scopes and prices per engagement and publishes no rate card.

Request an AI Assurance Review.

Request a confidential briefing