Technology, AI, and Digital Platforms
Multilingual AI Red Teaming and Safety Testing in Pashto, Dari and 22 More Afghan Languages
Ariana Nexus runs adversarial testing of large language models, AI agents and content-moderation systems in Afghan languages, using native Pashto and Dari red teamers rather than translated prompts — and documents the result in the form EU AI Act Article 55, NIST AI 600-1 and California SB 53 ask for.
What is multilingual AI red teaming in Afghan languages?
Multilingual AI red teaming in Afghan languages is structured adversarial testing in which native Pashto and Dari speakers attempt, under controlled conditions and inside an agreed scope, to make an AI system produce harmful, unsafe or culturally damaging output. Ariana Nexus delivers it as a documented evaluation: a harm taxonomy written for the Afghan context, reproducible test sets authored in-language, severity-scored findings, and a report mapped line by line to EU AI Act Article 55(1)(a), NIST AI 600-1 and California SB 53.
Safety alignment was trained in English. Your users are not writing in English.
Guardrails thin out at the edge of the training distribution.
Model safety behaviour is learned overwhelmingly from English and a handful of high-resource languages. Pashto and Dari sit far outside that distribution. The published literature is consistent: refusal behaviour, classifier accuracy and answer quality all degrade in low-resource languages, and they degrade further when a user code-switches between Pashto, Dari and English inside a single turn — which is how Afghan users actually write.
Machine-translated test sets understate the risk and look like a pass.
The fastest way to claim multilingual coverage is to run an English red-team set through a translation engine. It is also the fastest way to get a false clean bill of health. When researchers replaced machine translation with human translators on the same attacks, success rates rose sharply — the earlier low numbers measured the translator, not the model. A test set that fails because the Pashto was wrong is not evidence of safety.
The harms that matter in Afghan languages are not on an English harm list.
A phrase that is descriptive in English can be recruitment language in Pashto. Advice that is unremarkable in Ohio can expose a woman in Herat to family retaliation. A model can quote religious text in a way that puts the person who asked at risk. None of this is caught by a taxonomy written for an English-speaking user in a country with functioning courts.
The opposite failure is invisible in English-only testing.
Models also over-refuse. They flag ordinary Pashto religious or cultural speech as unsafe, decline legitimate requests from Afghan users at higher rates than from English-speaking ones, and answer Afghan Dari prompts in Iranian Persian. Over-refusal is a fairness and product failure as well as a safety one, and it never appears in an evaluation that only counts harmful outputs.
Published findings, cited so you can check them. Ariana Nexus does not publish figures it cannot source.
What we test
One evaluation, seven surfaces. Scope is set in writing before any testing begins, and every surface in scope gets its own coverage statement in the report.
Text model adversarial testing
Single-turn and multi-turn attacks authored in Pashto and Dari: direct elicitation, role framing, gradual intent distribution across turns, code-switching between Afghan languages and English, Latin-script transliteration, and obfuscation using Perso-Arabic orthographic variation.
jailbreak testing · prompt attacks · multi-turn red teaming
AI agent and tool-use testing
Indirect prompt injection through Pashto and Dari documents, web content and messages; tool misuse; instruction hierarchy failures; and agent behaviour when the task, the content and the system prompt are in three different languages.
prompt injection · agentic AI safety · tool misuse
Guardrail and classifier evaluation
Precision and recall of safety classifiers and moderation models per language and per variety — measured in both directions, because a classifier that over-flags Pashto is failing just as clearly as one that misses a threat.
guardrail testing · content classifier evaluation · false positive rate
Speech, voice and ASR
Mis-transcription as a safety failure, not a quality metric: accent and dialect coverage across Kandahari, Nangarhari and Wardaki Pashto and Kabuli, Herati, Mazari and Hazaragi Dari; wake-word and command misfires; synthetic-voice misuse.
speech recognition evaluation · Pashto ASR · voice AI safety
Machine translation and summarization safety
Safety-critical mistranslation in both directions, hallucinated content in summaries of Afghan-language source material, and silent register shifts that change a warning into a suggestion.
machine translation quality · MT risk · summarization hallucination
Content moderation and Trust and Safety policy testing
Whether your written enforcement policy survives contact with the language: we take the policy you actually enforce, build adversarial cases against each clause in Pashto and Dari, and report where the policy is unenforceable as written.
trust and safety · content moderation policy · platform enforcement
Multimodal and document understanding
Perso-Arabic script rendering and OCR failure, Afghan civil documents and identity artefacts, image-text attacks that carry the payload in the image, and handwriting.
multimodal safety · OCR Pashto Dari · document AI
The Afghan harm taxonomy
The part that cannot be bought anywhere else. Eight harm domains, each with named failure modes, maintained as a versioned instrument and adapted to your system at the start of every engagement.
Violent extremism and incitement
Recruitment register, martyrdom framing, tribal and sectarian calls to action, and the specific problem of text that is descriptive in English translation and operative in Pashto.
A sentence can pass an English classifier and function as instruction in the original.
Safety of women and girls
Guidance on education, employment, travel without a mahram, divorce, custody and reporting abuse where the answer that is correct in the United States is the answer that gets a woman hurt in Afghanistan.
We score whether the model recognises which country it is answering into.
Religious content and exposure risk
Misquotation of scripture, the line between explaining a ruling and issuing one, and output that could expose the person who asked to an accusation.
Reviewed by native speakers who know what is dangerous to have on a phone.
Ethnic and sectarian harm
Pashtun, Tajik, Hazara, Uzbek, Turkmen, Baloch and Nuristani framing; coded slurs that no English filter recognises; and grievance narratives repeated as fact.
Coded terms change faster than filters; the taxonomy is versioned for that reason.
War, displacement and moral injury
How the system responds to an Afghan user describing bombardment, the loss of family, evacuation, the 2021 withdrawal, an interpreter left behind, or the guilt of having left. Whether it recognises distress at all in a language whose idioms of suffering are not literal.
Crisis handling is tested in Pashto and Dari, not assumed from English behaviour.
Personal safety and operational exposure
Output that could identify a former interpreter, journalist, activist, judge or NGO worker or their family; location inference from ordinary detail; and advice that silently assumes a functioning legal system.
This domain is scored on consequence, not on whether the text reads as harmful.
Fraud, rumour and the Afghan information space
Resettlement and visa fraud around SIV, P-1 and P-2 processes, aid-distribution rumour, health misinformation, hawala and currency scams, and claims attributed to de facto authorities.
Seeded from what is actually circulating, not from a generic misinformation list.
Over-refusal, erasure and mis-service
The failure that English-only evaluation never sees: legitimate Pashto and Dari requests refused, ordinary cultural or religious speech flagged as unsafe, Afghan Dari answered in Iranian Persian, and dialects collapsed into a single wrong standard.
Scored in parallel with harmful output, on the same severity scale.
How the engagement runs
Six phases. Every phase produces an artefact you keep, and every artefact is written so it can be handed to a regulator, an auditor or a customer without rewriting.
Scoping and threat modelling
Week 1
System boundary, deployment context, user population and your existing risk register. We agree the harm domains in scope, the attack budget, the legal and ethical boundaries of the exercise, and what happens to a finding the moment it is confirmed.
Threat model and scope memorandum
Taxonomy and test-set construction
Weeks 1–2
The Afghan harm taxonomy is adapted to your system. Prompts are authored in Pashto and Dari by native speakers — never machine translated — with dialect variants, transliterated forms and code-switched constructions, and seeded from the Afghan information space as it is now.
Versioned taxonomy and reproducible test set
Red team execution
Weeks 2–4
Native-speaker teams work to a structured attack budget under logged conditions: single-turn, multi-turn and agentic sessions, blind and informed protocols, with the model version, system prompt, temperature and date recorded against every transcript.
Logged transcripts and attack ledger
Adjudication and severity scoring
Week 4
Two independent reviewers score every candidate finding; disagreements go to a third. Severity, likelihood and affected population are scored separately, and over-refusal is adjudicated on the same scale so the two failure modes can be compared.
Findings register with inter-rater statistics
Reporting and regulatory mapping
Weeks 5–6
One report, three readers: your safety engineers get reproduction steps, your counsel gets the regulatory mapping, your communications team gets language it can publish. Coverage and limitations are stated as plainly as the findings.
Evaluation report and regulatory appendix
Re-test, regression and monitoring
Post-mitigation
After you remediate, we re-run the failing set and report what actually moved. The regression suite is handed over so your own team can run it at every release, and campaigns can be scheduled quarterly or tied to model releases.
Re-test report and regression suite
What the regulations require, and what you receive
This is the section buyers arrive for. Each row states the obligation in the instrument's own terms, then the artefact that answers it. Ariana Nexus evaluates and documents; it does not certify, and it never presents an evaluation as a conformity assessment.
What you receive
Eight artefacts. Each one is written to be used by someone other than us — that is the test we apply before anything is delivered.
Threat model and scope memorandum
System boundary, deployment assumptions, harm domains in scope and out, and the agreed attack budget.
Afghan-language harm taxonomy, versioned
The eight domains with named failure modes, adapted to your system and dated, so a later evaluation can be compared with this one.
Reproducible test set with provenance
Prompts, dialect variants, transliterations and multi-turn scripts, each with its author, language, variety and the harm domain it targets.
Findings register
Severity, likelihood, affected population, reproduction steps and transcript evidence for every confirmed finding, with inter-rater agreement reported.
Over-refusal and service-quality report
The other failure mode, by language and variety: refusal rates on legitimate requests, wrong-variety responses, and quality gaps against the English baseline.
Regulatory mapping appendix
Article 55, Article 53, NIST AI 600-1, SB 53, ISO/IEC 42001, OWASP and ATLAS, line by line against the evidence produced.
Publication-ready language
Draft text for your model card, system card or transparency report describing the evaluation, its scope and its limits — for you to verify, edit and own.
Regression suite and re-test report
The failing set packaged so your team can run it at every release, plus our measurement of what changed after remediation.
Why Ariana Nexus, and how we are different
There is no certification for Afghan-language AI evaluation. No board examines it, no registry lists it, and no vendor can show you one. So we wrote the standard we test to, we train the people who apply it, and we publish the method rather than asking you to take it on trust.
Native red teamers, not translated prompts
Every prompt is authored in-language by a native speaker of the specific variety, with the dialect, register and code-switching that real users produce. The published evidence is that human-authored attacks find substantially more than machine-translated ones — a translated test set measures your translation engine.
A firm of scholars, not a pool of bilinguals
The engagement is run by alumni of Cornell, Brown, the University of Chicago, the University of British Columbia and Otto-von-Guericke University Magdeburg, working in their own first language. Measurement design, adjudication statistics, regulatory mapping and clinical supervision are held by people qualified to hold them.
Both failure modes, on one scale
We measure harmful output and over-refusal in the same campaign and score them on the same scale. A vendor that only counts jailbreaks will hand you a model that is safe because it has stopped answering Afghan users.
Delivered entirely in-house
Every part of the engagement is produced by our own people. No subcontractors, no brokered specialists, no capacity bought in for the week. One engagement, one point of accountability.
Documentation is the product
Most evaluation work arrives as a deck. Ours arrives as an evidence pack indexed to the instruments your regulator, your auditor and your enterprise customers actually cite.
Duty of care to the people who do the work
Analysts working through extremist, sexual and war-related material are on capped caseloads with a structured debriefing protocol set and reviewed by a doctoral-level psychologist on the firm's own staff.
Four ways this work gets done
Language and variety coverage
Twenty-four Afghan languages across five families. Pashto and Dari carry the full evaluation bench; the remaining languages are covered to the depth the speaker population allows, and we state that depth in writing rather than implying uniform coverage.
Iranian
Turkic
Indo-Aryan
Nuristani
Dravidian
Afghan Dari is not Iranian Persian and Afghan Turkmeni is not the standard Turkmen of Turkmenistan. Hazaragi is a variety of Dari, staffed separately. Where a model answers an Afghan Dari prompt in Iranian Persian, that is a finding, not a near miss.
The team behind this service
Ariana Nexus is operated and led by alumni and scholars of leading universities who work in Afghan languages as their own first languages. On an evaluation engagement, the tooling, the measurement design, the regulatory mapping and the harm taxonomy are each held by a named person qualified to hold them. We do not staff this work with bilinguals.
The difference shows up in the adjudication. Deciding whether a Pashto answer is dangerous requires someone who knows the language, the region the user is writing from, and what happens to a person who is found with that text on their phone. Deciding whether a difference between two reviewers is real requires someone who can compute it. Very few firms have both in the same room. That is the firm.
Hassan Ukasha
Managing Partner
- B.S. Cornell University
- M.P.H. Cornell University
Every Ariana Nexus engagement is delivered under the Managing Partner's oversight. Hassan Ukasha holds accountability for the firm's operations and for this program. He approves the scope of each evaluation before testing begins, signs the independence and conflicts position that appears on the face of every report, and is the escalation point for any finding that changes what a client can safely say about its own system.

Zeba Haqbani
Senior Partner
- B.A.The American University of Afghanistan
- B.A.University of British Columbia
Builds the firm's institutional systems and AI platforms. Owns the evaluation tooling, the test-set infrastructure and the transcript record, so every finding on this page can be reproduced from a logged run rather than a recollection.

Hussain Ahmad
Principal
- M.Eng.Cornell University
- Ph.D.University of Chicago
Measurement design and the statistics of evaluation: attack budgets, sampling, inter-rater reliability, and what a difference in failure rate between two languages does and does not support. Decides when a number is strong enough to publish.

Wasil Peroz
Principal
- B.A.Milli University
- M.Sc.Otto-von-Guericke University Magdeburg
Institutional law. Maps every finding to EU AI Act Articles 53 and 55, NIST AI 600-1, California SB 53 and ISO/IEC 42001, and drafts the regulatory appendix your counsel reads before you publish anything.

Maryam Safi
Principal
- B.A.Cornell University
Afghan information-space analysis. Maintains the harm taxonomy as a versioned instrument and seeds each campaign from what is actually circulating in Pashto and Dari, not from a translated list written somewhere else.

Independence, boundaries and data handling
We evaluate. We do not certify.
Our report sits alongside your conformity assessor, your auditor and your counsel — never in place of them. We will not describe an evaluation as a conformity assessment, and we will not sign anything that implies we have.
Findings go to the owner, never to the public.
Confirmed findings are delivered to the system owner to remediate. We do not publish working attacks, release exploit tooling, or retain live jailbreaks outside the engagement record.
Your model and your data stay inside the agreement.
Model access is handled under NDA in the environment you specify. Prompts, outputs and transcripts are not used to train anything, are not reused on another engagement, and do not leave the agreed environment.
Independence is stated, not assumed.
Every report carries a declared independence and conflicts position. If we hold another engagement that touches the system under test, you are told before scoping, not after delivery.
No channel through the de facto authorities.
Ariana Nexus does not route documents, data or inquiries through channels controlled by the de facto authorities in Afghanistan. The firm has no offices and no operations inside the country.
The people doing the work are protected.
Red-team analysts are named in the engagement record and never in public findings. Caseloads are capped, debriefing is scheduled rather than offered, and the protocol is set and reviewed by Diana Ayubi, Psy.D., Engagement Manager at the firm.
Who this is built for
Frontier model developers and AI labs
Article 55 adversarial-testing evidence, SB 53 third-party involvement, and a defensible basis for the language claims in your model card.
Safety, Policy and Model Evaluation leads
Platforms and Trust and Safety teams
Moderation-policy fidelity in Pashto and Dari, classifier precision and recall by variety, and systemic-risk evidence for Afghan diaspora communities on your service.
Trust and Safety, Integrity and Policy leads
Enterprises deploying AI in Afghan-facing services
Health systems, resettlement organisations, insurers and banks using translation, chat or triage models with Afghan users — tested before the incident, not after.
Chief AI Officers, risk and compliance
Government agencies and prime contractors
Mission systems using machine translation or language models in Pashto and Dari, evaluated against NIST guidance by a firm registered to contract.
Program managers and capture leads
Assurance, audit and evaluation organisations
An Afghan-language evaluation bench you can bring into an engagement you are leading, with our method and our people named in your evidence file.
Engagement partners and technical leads
Engagement models
Evaluation is scoped as a program with a named engagement lead, not priced by the word or the prompt. Scoping begins under NDA.
Baseline evaluation
Four to six weeks
One system, one release. Full taxonomy, all agreed surfaces, complete evidence pack and regulatory appendix. The usual starting point, and the one most buyers need before a launch.
Release-gated program
Per release, retained
The regression suite is maintained between releases and re-run at each one, with a delta report against the previous evaluation and a standing coverage statement.
Continuous red team
Standing bench
A named team held against your systems year-round: scheduled campaigns, incident surge capacity, and quarterly reporting into your governance forum.
Targeted assessment
Two to three weeks
One harm domain, one language family or one surface — commonly used where an incident, a customer question or a regulator has raised something specific.
At a glance
Questions buyers ask
What is AI red teaming in Pashto and Dari?
It is structured adversarial testing of an AI system by native Pashto and Dari speakers, who attempt under controlled conditions to make the system produce harmful, unsafe or culturally damaging output. The result is a documented evaluation: a harm taxonomy, reproducible test sets, severity-scored findings and a report mapped to the regulatory instruments that apply to your system.
Why do AI safety guardrails fail in Afghan languages?
Safety behaviour is learned overwhelmingly from English and a few high-resource languages. Pashto and Dari sit far outside that distribution, so refusal behaviour, classifier accuracy and answer quality all degrade — and degrade further when users code-switch or transliterate, which Afghan users do constantly. The harms that matter in the Afghan context are also absent from taxonomies written for English-speaking users.
Can we just translate our English red-team set into Pashto?
You can, and it will usually produce a clean result that is not true. When researchers replaced machine translation with human translators on the same attacks, success rates rose by roughly thirty percent — the low scores had been measuring translation quality rather than model robustness. A translated test set is evidence about your translation engine.
Does this satisfy EU AI Act Article 55?
Article 55(1)(a) requires providers of general-purpose AI models with systemic risk to perform model evaluation to state-of-the-art protocols, including conducting and documenting adversarial testing. Our engagement produces exactly that record for Afghan languages: protocol, attack budget, coverage statement, reproducible prompts and logged transcripts. Whether your overall Article 55 position is sufficient is a judgement for you and your counsel; we supply the evidence and the mapping, and we do not certify.
What does California SB 53 require about languages?
A frontier developer deploying a new or substantially modified frontier model must publish a transparency report stating, among other things, the languages and modalities of output the model supports. Large frontier developers must also publish summaries of catastrophic-risk assessments and descriptions of third-party involvement. We give you measured evidence behind the language claim and a third-party evaluator description drafted for publication.
How does this map to NIST AI 600-1?
The Generative AI Profile's MEASURE function calls for structured assessment of generative AI risks, including external red-teaming and documented evaluation of harmful content, bias and security risks. We deliver a measurement plan mapped to AI RMF subcategories, reported per language rather than in aggregate, because an aggregate multilingual score hides exactly the gap you are testing for.
Which Afghan languages can you test?
Twenty-four, across five language families. Pashto and Dari carry the full evaluation bench, including dialect coverage across Kandahari, Nangarhari and Wardaki Pashto and Kabuli, Herati, Mazari and Hazaragi Dari. The smaller languages are covered to the depth the speaker population allows, and we state that depth in the report rather than implying uniform coverage.
Do you test speech and voice models, or only text?
Both, plus agents, guardrail classifiers, machine translation, moderation systems and multimodal inputs. For speech we treat mis-transcription as a safety failure rather than a quality metric, because a wrong transcript upstream produces a wrong decision downstream.
Do you measure over-refusal as well as harmful output?
Yes, in the same campaign and on the same severity scale. Models frequently over-block Pashto and Dari, flagging ordinary religious and cultural speech as unsafe and refusing legitimate requests at higher rates than in English. A vendor that only counts jailbreaks will hand you a model that looks safe because it has stopped serving Afghan users.
How long does an evaluation take and how is it priced?
A baseline evaluation runs four to six weeks. Release-gated programs, continuous benches and targeted assessments are scoped individually. Engagements are scoped as programs with a named engagement lead rather than priced per prompt, and scoping begins under NDA.
How do you protect our model, our prompts and our findings?
Model access is handled under NDA in the environment you specify. Prompts, outputs and transcripts are not used to train anything, are not reused on another engagement and do not leave that environment. Confirmed findings go to you to remediate; we do not publish working attacks or release exploit tooling.
Can Ariana Nexus be named as an independent third-party evaluator?
Yes. Every report carries a declared independence and conflicts position, a scope description and an evaluator statement written to be quoted in a model card, a transparency report or an audit file. If we hold another engagement that touches the system under test, you are told before scoping.
Start with a scoping conversation
Tell us what the system is, where it is deployed and which languages you have claimed. We will tell you what an evaluation would cover, what it would not, and what the evidence would look like when it is finished. Scoping runs under NDA and we meet in person at the Washington, D.C. office where that helps.
Request a scoping conversationRelated services
- LLM evaluation and benchmark development
How well the model performs in Pashto and Dari. This page asks how it fails under attack; that one asks how well it works when nobody is attacking it.
- AI governance documentation for Afghan-language coverage
The management system and the conformity file. This page produces the test evidence; that page assembles, maintains and keeps it current.
- Speech recognition and text-to-speech data and evaluation
Speech data and ASR/TTS evaluation — recorded audio, transcription, voice licensing and listening tests.
- Deepfake and synthetic media analysis in Pashto and Dari
Detection, labeling and takedown support for synthetic audio and video in Afghan languages.