-->

LLM Evaluation & Red-Teaming in Afghan Languages

Your model answers in fluent Pashto. That proves it is confident — not that it is correct, and not that anyone has checked.

Independent validation, multilingual evaluation, and red-teaming of AI systems across all 24 Afghan languages — where models are weakest, least tested, and most confidently wrong. Aligned to the NIST AI Risk Management Framework, with assurance performed by native-speaker experts who can actually read the output.

Convened by Ariana Nexus · AI & Data Systems Practice · Washington, D.C.

Request an AI Assurance Review
24Afghan languages
0Security incidents
100%Senior-led engagements
NIST AI RMFAligned · independent
The Problem

A model does not warn you when it is wrong in Dari

A model does not warn you when it is wrong in Dari. It gives you the same confident answer it gives in English.

A large language model produces fluent output in every language it touches. But fluency in a low-resource language is not accuracy — and the model offers no signal to tell the two apart. It answers a Pashto prompt with the same confidence whether it is right or catastrophically wrong, and the English-language benchmarks it passed say nothing about either.

The failure is structural. Models are trained on the least data, tested with the fewest benchmarks, and guarded by the weakest safety filters in exactly the low-resource languages where confidence outruns competence. A safety guardrail that holds in English is routinely bypassed by the same request in a language no red team evaluated. The hallucination, the unsafe output, the cultural error — none of it surfaces in evaluation, because the people evaluating cannot read the language.

Governance has caught up to the obligation even as the politics shift around it. The NIST AI Risk Management Framework asks every deployer to measure and manage AI risk; the EU AI Act attaches conformity obligations to high-risk systems; state laws are following. None of that is satisfiable by an English benchmark for a model that operates in 24 languages.

Ariana Nexus validates and red-teams AI systems in the languages they actually serve — fluency separated from accuracy, guardrails tested where they are weakest, and every finding documented for governance and disclosed for repair.

FluencyAccuracyHigh-resourceLow-resourcedivergence
The gap no benchmark sees — fluent output, in a language the evaluation never tested.
24Afghan languages evaluated — most of them low-resource, where models fail and no one checks.
NIST AI RMFMeasure and Manage, in the languages your benchmarks skip.
Sovereign Speech IndexOur model-performance measure.

Fluency is not accuracy.

Evidence Ledger

What the evidence shows — and your benchmark cannot

Every figure below is published and independently sourced. None of it is visible to an English-language evaluation.

FindingWhat it meansSource
79%Unsafe prompts translated into low-resource languages bypassed a frontier model's guardrails 79% of the time — versus under 1% in English. Safety training does not transfer across languages.Yong, Menghini and Bach, 2023–24
~35%Roughly 35% of responses to malicious prompts in low-resource languages contained harmful content — versus about 1% in high-resource languages.Shen et al., ACL 2024
44%Even best-case English clinical summarization hallucinated — and 44% of those hallucinations were clinically significant.Asgari et al., npj Digital Medicine, 2025
12–20 ptsThe best frontier model scored 12 to 20 accuracy points lower in lower-resource languages than in English.AAAI 2025
<45%For the most under-resourced languages — the category Pashto and Dari fall into — most frontier models score below 45% accuracy.IndicParam, 2026
49.1%When an adverse event occurs, Limited-English-Proficiency patients are harmed in 49.1% of cases — versus 29.5% for other patients.Divi et al., 2007
25.9MAbout 25.9 million people in the U.S. (~9%) have Limited English Proficiency and are owed language access under federal law.U.S. Census ACS / MPI

Figures are published, peer-reviewed where applicable, and current as of June 2026. Full citations available on request.

Definition

What is AI model validation and red-teaming for low-resource languages?

Model Validation & Red-Teaming is independent evaluation, validation, and adversarial red-teaming of AI and large language model systems for their behavior in Afghan languages — across all 24, most of them low-resource, where models are weakest and least tested. It is aligned to the NIST AI Risk Management Framework and its Generative AI Profile, and covers multilingual accuracy evaluation, the Cultural Hallucination Audit for fluent-but-wrong output, safety-guardrail and jailbreak red-teaming, and bias testing — performed by native-speaker subject-matter experts and documented for AI governance. Ariana Nexus reports how a model actually behaves in the languages it claims to serve; red-teaming is authorized and responsibly disclosed, with findings sent to the owner to remediate rather than weaponize.

In brief

Large language models are trained on the least data and guarded by the weakest safety filters in low-resource languages, so they answer fluently but are more often wrong — and standard English benchmarks never detect it. Validating them requires native-speaker evaluation, hallucination auditing, and red-teaming in the target language, documented to the NIST AI Risk Management Framework.

Operating Model

One practice. Three coordinated capabilities.

Three institutional capabilities, orchestrated into assurance you can put in front of an auditor.

HIC
Human Intelligence Collective

Lived-expertise practitioners across all 24 Afghan languages — the cultural gatekeepers who keep every engagement anchored in ground truth, never extractive.

Native-speaker subject-matter experts across all 24 languages who read model output critically, catch culturally coded errors and hallucinations, and serve as expert human evaluators and red-teamers — the ground truth no automated metric replaces.

Protocol — The Cultural Hallucination Audit
ADF
AI Data Factory

Governed Afghan-language data infrastructure, evaluation benchmarks, and institutional-grade training assets meeting auditable standards.

Afghan-language evaluation datasets and benchmarks, the Cultural Hallucination Audit pipeline, multilingual evaluation harnesses, adversarial and red-team prompt sets, and human-in-the-loop scoring.

Protocol — The ADF Pipeline
CCB
Cultural Compliance Bureau

An audit-grade review regime translating cultural intelligence into compliance-ready practice — the governance layer threading through every engagement.

Methodology rigor and independence audit; reproducibility; bias and fairness review; responsible-disclosure governance for red-team findings; the CCB Sign-Off Mark on every assurance report.

Protocol — The CCB Sign-Off Mark
HICADFCCBOne governed assurance report

Three capabilities. One honest answer about what your model does.

The Path · Cultural Hallucination Audit

How Ariana Nexus exposes fluent-but-wrong output: the Cultural Hallucination Audit

Integrated 4-phase system · 3 institutional capabilities · 5 validation gates

The Cultural Hallucination Audit™ exposes fluent-but-wrong output; the Five-Gate Validation Protocol™ governs the whole assurance engagement, evaluation to red-team to sign-off.

The Five Gates

1Linguistic2Cultural3Standards4Population5Sign-Off
01

Linguistic Accuracy

Model outputs evaluated for linguistic accuracy across 24 languages by native-speaker experts; the Pashto–Dari Parity Index applied; automated metrics validated against human judgment, never trusted alone.

02

Cultural Validity

Outputs evaluated for cultural validity and culturally coded error; the Cultural Hallucination Audit applied; cleared by the CCB Sign-Off Mark.

03

Standards Conformance

NIST AI Risk Management Framework (Govern, Map, Measure, Manage) and the Generative AI Profile; EU AI Act conformity where applicable; recognized AI red-teaming and evaluation practice.

04

Population Risk

Safety-guardrail robustness across languages, including low-resource jailbreak resistance; bias and harmful-output testing; responsible disclosure of vulnerabilities to the owner — no weaponization, no public exploit tooling.

05

Institutional Sign-Off

Evaluation and red-team findings documented, reproducible, traceable, and ready for AI governance and audit.

The Four-Phase Orchestration Cycle

IIIIIIIV
I

Situation — Understand

The system, its intended use and languages, and the assurance requirements — NIST AI RMF, EU AI Act, internal — mapped.

Cultural mapping · stakeholder calibration · constraint discovery
II

Complication — Architect

The evaluation and red-team plan, benchmarks, and Cultural Hallucination Audit designed for the target languages.

Program scaffolding · compliance baseline · governance charter
III

Resolution — Deploy

Multilingual evaluation and red-teaming executed; failures detected, characterized, and reported for remediation.

In-context execution · data infrastructure
IV

Measured Outcome — Govern

Findings documented for governance; re-tested after remediation; benchmarked on the Sovereign Speech Index and the Low-Resource Safety Benchmark; monitored across the model lifecycle.

Continuous documentation · red-team validation · multi-decade horizon

Active throughout: HIC and CCB at full intensity on every evaluation; ADF heaviest at Phases II–III.

Native-speaker assurance

Evaluation performed by people who can actually read the output — the ground truth no automated metric replaces.

Mandate Register

The obligations your validation evidence has to satisfy

A living register of the frameworks and rules that govern AI validation and red-teaming — with current status, because this audience is not served by stale regulatory framing.

Regulation / StandardWhat it requiresStatus
NIST AI RMF 1.0Govern, Map, Measure, Manage across the AI lifecycle; red-teaming sits under Measure.In effectVoluntary · Jan 2023
NIST AI 600-1 — GenAI ProfileTwelve generative-AI risk categories that scope a red-team test plan.FinalJul 2024
EU AI Act — GPAI obligationsModel evaluation, adversarial testing and red-teaming, and incident reporting for systemic-risk models.ApplicableAug 2025 · enforceable Aug 2026
EU AI Act — High-risk (Annex III)Accuracy, robustness, data governance, human oversight, and conformity assessment.Applies Aug 2026Digital Omnibus may defer to Dec 2027
ISO/IEC 42001:2023Certifiable AI management system (AIMS) — the auditable governance shell.In effectCertifiable · 2023
ISO/IEC 23894:2023AI risk-management methodology adapting ISO 31000 to AI.In effectGuidance · 2023
FDA — PCCP for AI-enabled devicesPre-specified protocol governing how AI/ML models may be updated after clearance.Final guidanceDec 2024
FDA — AI device lifecycle managementTPLC validation, transparency, bias management, and post-market monitoring.Draft — not bindingJan 2025
ONC HTI-1 — Decision SupportSource-attribute transparency and intervention risk management, including fairness, for certified EHRs.In effectDSI compliance Dec 2024
HHS Section 1557 (2024)Affirmative duty to identify and mitigate discrimination in clinical decision-support algorithms.In effectCompliance May 2025 · partial vacatur Oct 2025
OWASP Top 10 for LLM AppsThe de facto checklist for LLM application red-teaming and security testing.Current2025 edition
MITRE ATLASAdversarial-ML tactics and techniques taxonomy structuring red-team exercises.Livingv2026.05
TJC + CHAI — Responsible Use of AIGovernance, bias assessment, and continuous monitoring expectations for health systems.VoluntarySep 2025

Status current as of June 2026. Draft and voluntary instruments are marked as such; we never present a draft as a binding requirement.

Without In-Language Validation

What happens without in-language validation

AI systems deployed into Afghan-language use without in-language validation did not announce their failures. The model answered fluently in Pashto — and was confidently wrong often enough to mislead — while every English-language benchmark it passed reported it as fine. Safety guardrails that held in English were bypassed in a low-resource language no red team had tested.

The hallucination, the unsafe output, the cultural error: none surfaced in evaluation, because no one evaluating could read the language. They surfaced in production, in front of the people the system was supposed to serve — and by then they were not findings in a report but harms in the world.

An unvalidated model in Pashto is an untested model in production.

What Partnership Looks Like

Your model, tested in the languages it serves

From foundations to continuous stewardship.

1 / 4
Foundations

Scoped, mapped, architected. The system, its languages, and its assurance requirements understood.

2 / 4
Activation

Built to standard. The evaluation and red-team plan, benchmarks, and Cultural Hallucination Audit designed for the target languages.

3 / 4
Operating Rhythm

The active state. Evaluation and red-teaming running; failures found, characterized, and disclosed for repair.

4 / 4
Continuous Stewardship

Across the model lifecycle. Re-tested after remediation; benchmarked; monitored as the model and its risks evolve.

The Receivables

model opacityauditable truth
Independent multilingual evaluation across 24 Afghan languages.Accuracy, fluency, and reliability scored by native-speaker experts — not automated metrics alone.
A Cultural Hallucination Audit.Fluent-but-wrong output and culturally coded errors detected, where standard evaluation sees nothing.
AI red-teaming for low-resource languages.Safety-guardrail robustness, jailbreak resistance, bias, and harmful-output testing — responsibly disclosed for remediation.
NIST AI RMF-aligned assurance documentation.Measure-and-Manage evidence your governance and auditors can use.
A Sovereign Speech Index score for your model.Performance across the 24 languages, benchmarked.
A Low-Resource Safety Benchmark result.Where the guardrails hold, and where they do not.
Reproducible methods and documented independence.Findings another evaluator can verify.
Re-testing after remediation, and lifecycle monitoring.Assurance that continues as the model evolves — not a one-time certificate.

What you receive is not a green checkmark from an English benchmark. It is the truth about how your model behaves in the languages it claims to serve.

Readiness Ladder

Where your AI assurance actually stands

Five levels of validation maturity for multilingual AI. Most teams overestimate where they are.

L0
Self-attested
Vendor claims and English benchmarks. No independent evidence in the languages served.
L1
English-validated
Tested in English; low-resource-language behavior is unknown and unmeasured.
L2
Spot-checked Most teams sit here
Ad hoc in-language review, not reproducible and not documented for governance.
L3
Independently validated
Native-speaker evaluation and red-teaming, documented to the NIST AI Risk Management Framework.
L4
Continuously assured
Re-tested after remediation, benchmarked, and monitored across the model lifecycle. Audit-ready.

Ariana Nexus engagements move a model from L0–L2 to L3 and hold it at L4.

Global Coverage

The framework changes. A model wrong in Pashto is wrong everywhere.

AI assurance obligations span the United States, the EU AI Act’s conformity regime, the United Kingdom’s evaluation work, and enterprise and government deployments worldwide that touch Afghan languages. The regime differs by jurisdiction; the failure mode does not. Ariana Nexus validates and red-teams AI systems in all 24 Afghan languages, worldwide.

EUROPEGULF & ARAB WORLDUnited KingdomGermanyNetherlandsFranceItalySwedenUnited Arab EmiratesSaudi ArabiaQatarUnited StatesPRIMARY ANCHOR

— and worldwide.

The regulation changes. The model still has to be right in the language it answers.

Who Leads

Who leads the AI & Data Systems Practice

Hussain Ahmad, Practice Leader, AI Engineering, Ariana Nexus
Hussain Ahmad
Practice Leader, AI Engineering

Model evaluation, multilingual AI systems, and assurance engineering · Cornell · University of Chicago.

Maryam Safi, Principal, Cultural Compliance Bureau, Ariana Nexus
Maryam Safi
Principal, Cultural Compliance Bureau (CCB)

Cultural compliance and responsible disclosure · Cornell University.

Every engagement is senior-led.

The Door

Request an AI Assurance Review.

For AI developers and model providers, government AI programs, enterprises deploying AI in Afghan-language contexts, and AI-governance and assurance teams. Independent and reproducible. Briefings are conducted under NDA, in Washington, D.C. or virtually.

Request a confidential briefing

Have a specific question, a use case, or a considered perspective on our work? We welcome substantive input from the institutions we serve. Share it with us.

Until someone can read the output, the model is not validated — only fluent.

The Cultural Hallucination Audit™ · Standards adherence (NIST AI RMF, the Generative AI Profile, the EU AI Act) · Five-Gate Validation Protocol™ · The Sovereign Speech Index · The Low-Resource Safety Benchmark · Responsible-disclosure policy. Full index at /assurance/.

Offices: Washington, D.C. · Arlington, VA · London (planned) · Berlin (planned) · NAICS 541930 · SAM.gov UEI M2UDMUDFXGL9 · CAGE 1Z3A3 · Section 1557-ready · NIST AI RMF-aligned · ISO/IEC 27001 · GDPR / UK GDPR · FedRAMP-aligned · Worldwide service.