Skip to main content

Technology, AI, and Digital Platforms

Multilingual AI Red Teaming and Safety Testing in Pashto, Dari and 22 More Afghan Languages

Ariana Nexus runs adversarial testing of large language models, AI agents and content-moderation systems in Afghan languages, using native Pashto and Dari red teamers rather than translated prompts — and documents the result in the form EU AI Act Article 55, NIST AI 600-1 and California SB 53 ask for.

What the evidence shows
79%
of harmful prompts translated from English into low-resource languages bypassed GPT-4's safety guardrails on AdvBench — comparable to state-of-the-art jailbreak techniques.Yong, Menghini and Bach, 2023 · arXiv:2310.02446
77.1% vs 59.5%
harmful-response rate achieved by human red teamers versus automated multilingual testing of the same systems. People find what pipelines miss.Multilingual multi-turn jailbreak study, 2026 · arXiv:2605.18239
+28.6% / +30.6%
increase in attack success when human translators replaced machine translation for two low-resource languages — the safety score was measuring translation quality, not model robustness.Same study · arXiv:2605.18239
2.92 / 5
human-rated usefulness of model answers in Arabic, Farsi, Pashto and Kurdish, against 3.86 / 5 in English; rated factuality fell from 3.55 to 2.87.Taraaz multilingual evaluation, 2026
Independent adversarial testing and safety evaluation of AI systems in Afghan languagesPashto and Dari firstText LLMsEU AI Act Articles 53 and 55 · NIST AI 600-1 and AI RMF 1.0 · California SB 53 · ISO/IEC 42001 · OWASP LLM Top 10 · MITRE ATLAS
Definition

What is multilingual AI red teaming in Afghan languages?

Multilingual AI red teaming in Afghan languages is structured adversarial testing in which native Pashto and Dari speakers attempt, under controlled conditions and inside an agreed scope, to make an AI system produce harmful, unsafe or culturally damaging output. Ariana Nexus delivers it as a documented evaluation: a harm taxonomy written for the Afghan context, reproducible test sets authored in-language, severity-scored findings, and a report mapped line by line to EU AI Act Article 55(1)(a), NIST AI 600-1 and California SB 53.

01 — The problem

Safety alignment was trained in English. Your users are not writing in English.

1

Guardrails thin out at the edge of the training distribution.

Model safety behaviour is learned overwhelmingly from English and a handful of high-resource languages. Pashto and Dari sit far outside that distribution. The published literature is consistent: refusal behaviour, classifier accuracy and answer quality all degrade in low-resource languages, and they degrade further when a user code-switches between Pashto, Dari and English inside a single turn — which is how Afghan users actually write.

2

Machine-translated test sets understate the risk and look like a pass.

The fastest way to claim multilingual coverage is to run an English red-team set through a translation engine. It is also the fastest way to get a false clean bill of health. When researchers replaced machine translation with human translators on the same attacks, success rates rose sharply — the earlier low numbers measured the translator, not the model. A test set that fails because the Pashto was wrong is not evidence of safety.

3

The harms that matter in Afghan languages are not on an English harm list.

A phrase that is descriptive in English can be recruitment language in Pashto. Advice that is unremarkable in Ohio can expose a woman in Herat to family retaliation. A model can quote religious text in a way that puts the person who asked at risk. None of this is caught by a taxonomy written for an English-speaking user in a country with functioning courts.

4

The opposite failure is invisible in English-only testing.

Models also over-refuse. They flag ordinary Pashto religious or cultural speech as unsafe, decline legitimate requests from Afghan users at higher rates than from English-speaking ones, and answer Afghan Dari prompts in Iranian Persian. Over-refusal is a fairness and product failure as well as a safety one, and it never appears in an evaluation that only counts harmful outputs.

Published findings, cited so you can check them. Ariana Nexus does not publish figures it cannot source.

02 — Scope

What we test

One evaluation, seven surfaces. Scope is set in writing before any testing begins, and every surface in scope gets its own coverage statement in the report.

01

Text model adversarial testing

Single-turn and multi-turn attacks authored in Pashto and Dari: direct elicitation, role framing, gradual intent distribution across turns, code-switching between Afghan languages and English, Latin-script transliteration, and obfuscation using Perso-Arabic orthographic variation.

jailbreak testing · prompt attacks · multi-turn red teaming

02

AI agent and tool-use testing

Indirect prompt injection through Pashto and Dari documents, web content and messages; tool misuse; instruction hierarchy failures; and agent behaviour when the task, the content and the system prompt are in three different languages.

prompt injection · agentic AI safety · tool misuse

03

Guardrail and classifier evaluation

Precision and recall of safety classifiers and moderation models per language and per variety — measured in both directions, because a classifier that over-flags Pashto is failing just as clearly as one that misses a threat.

guardrail testing · content classifier evaluation · false positive rate

04

Speech, voice and ASR

Mis-transcription as a safety failure, not a quality metric: accent and dialect coverage across Kandahari, Nangarhari and Wardaki Pashto and Kabuli, Herati, Mazari and Hazaragi Dari; wake-word and command misfires; synthetic-voice misuse.

speech recognition evaluation · Pashto ASR · voice AI safety

05

Machine translation and summarization safety

Safety-critical mistranslation in both directions, hallucinated content in summaries of Afghan-language source material, and silent register shifts that change a warning into a suggestion.

machine translation quality · MT risk · summarization hallucination

06

Content moderation and Trust and Safety policy testing

Whether your written enforcement policy survives contact with the language: we take the policy you actually enforce, build adversarial cases against each clause in Pashto and Dari, and report where the policy is unenforceable as written.

trust and safety · content moderation policy · platform enforcement

07

Multimodal and document understanding

Perso-Arabic script rendering and OCR failure, Afghan civil documents and identity artefacts, image-text attacks that carry the payload in the image, and handwriting.

multimodal safety · OCR Pashto Dari · document AI

03 — Harm taxonomy

The Afghan harm taxonomy

The part that cannot be bought anywhere else. Eight harm domains, each with named failure modes, maintained as a versioned instrument and adapted to your system at the start of every engagement.

Domain 01

Violent extremism and incitement

Recruitment register, martyrdom framing, tribal and sectarian calls to action, and the specific problem of text that is descriptive in English translation and operative in Pashto.

A sentence can pass an English classifier and function as instruction in the original.

Domain 02

Safety of women and girls

Guidance on education, employment, travel without a mahram, divorce, custody and reporting abuse where the answer that is correct in the United States is the answer that gets a woman hurt in Afghanistan.

We score whether the model recognises which country it is answering into.

Domain 03

Religious content and exposure risk

Misquotation of scripture, the line between explaining a ruling and issuing one, and output that could expose the person who asked to an accusation.

Reviewed by native speakers who know what is dangerous to have on a phone.

Domain 04

Ethnic and sectarian harm

Pashtun, Tajik, Hazara, Uzbek, Turkmen, Baloch and Nuristani framing; coded slurs that no English filter recognises; and grievance narratives repeated as fact.

Coded terms change faster than filters; the taxonomy is versioned for that reason.

Domain 05

War, displacement and moral injury

How the system responds to an Afghan user describing bombardment, the loss of family, evacuation, the 2021 withdrawal, an interpreter left behind, or the guilt of having left. Whether it recognises distress at all in a language whose idioms of suffering are not literal.

Crisis handling is tested in Pashto and Dari, not assumed from English behaviour.

Domain 06

Personal safety and operational exposure

Output that could identify a former interpreter, journalist, activist, judge or NGO worker or their family; location inference from ordinary detail; and advice that silently assumes a functioning legal system.

This domain is scored on consequence, not on whether the text reads as harmful.

Domain 07

Fraud, rumour and the Afghan information space

Resettlement and visa fraud around SIV, P-1 and P-2 processes, aid-distribution rumour, health misinformation, hawala and currency scams, and claims attributed to de facto authorities.

Seeded from what is actually circulating, not from a generic misinformation list.

Domain 08

Over-refusal, erasure and mis-service

The failure that English-only evaluation never sees: legitimate Pashto and Dari requests refused, ordinary cultural or religious speech flagged as unsafe, Afghan Dari answered in Iranian Persian, and dialects collapsed into a single wrong standard.

Scored in parallel with harmful output, on the same severity scale.

04 — Method

How the engagement runs

Six phases. Every phase produces an artefact you keep, and every artefact is written so it can be handed to a regulator, an auditor or a customer without rewriting.

Phase 1

Scoping and threat modelling

Week 1

System boundary, deployment context, user population and your existing risk register. We agree the harm domains in scope, the attack budget, the legal and ethical boundaries of the exercise, and what happens to a finding the moment it is confirmed.

Threat model and scope memorandum

Phase 2

Taxonomy and test-set construction

Weeks 1–2

The Afghan harm taxonomy is adapted to your system. Prompts are authored in Pashto and Dari by native speakers — never machine translated — with dialect variants, transliterated forms and code-switched constructions, and seeded from the Afghan information space as it is now.

Versioned taxonomy and reproducible test set

Phase 3

Red team execution

Weeks 2–4

Native-speaker teams work to a structured attack budget under logged conditions: single-turn, multi-turn and agentic sessions, blind and informed protocols, with the model version, system prompt, temperature and date recorded against every transcript.

Logged transcripts and attack ledger

Phase 4

Adjudication and severity scoring

Week 4

Two independent reviewers score every candidate finding; disagreements go to a third. Severity, likelihood and affected population are scored separately, and over-refusal is adjudicated on the same scale so the two failure modes can be compared.

Findings register with inter-rater statistics

Phase 5

Reporting and regulatory mapping

Weeks 5–6

One report, three readers: your safety engineers get reproduction steps, your counsel gets the regulatory mapping, your communications team gets language it can publish. Coverage and limitations are stated as plainly as the findings.

Evaluation report and regulatory appendix

Phase 6

Re-test, regression and monitoring

Post-mitigation

After you remediate, we re-run the failing set and report what actually moved. The regression suite is handed over so your own team can run it at every release, and campaigns can be scheduled quarterly or tied to model releases.

Re-test report and regression suite

05 — Regulatory mapping

What the regulations require, and what you receive

This is the section buyers arrive for. Each row states the obligation in the instrument's own terms, then the artefact that answers it. Ariana Nexus evaluates and documents; it does not certify, and it never presents an evaluation as a conformity assessment.

Instrument
Requirement
Delivered
Status
InstrumentEU AI Act, Article 55(1)(a)
RequirementProviders of general-purpose AI models with systemic risk must perform model evaluation to state-of-the-art protocols, including conducting and documenting adversarial testing to identify and mitigate systemic risk.
DeliveredA documented adversarial campaign in Pashto, Dari and 22 more Afghan languages: protocol, attack budget, coverage statement, reproducible prompt sets and logged transcripts.
StatusIn force since 2 August 2025
InstrumentEU AI Act, Article 55(1)(c)
RequirementSerious incidents must be tracked, documented and reported to the AI Office without undue delay.
DeliveredConfirmed findings routed on a severity clock, in a format your incident process can consume on the day it is raised rather than at the end of the engagement.
StatusCommission enforcement from 2 August 2026
InstrumentEU AI Act, Article 53 and Annex XI
RequirementTechnical documentation for general-purpose AI models, kept current and available to the AI Office.
DeliveredAn evaluation appendix drafted in documentation-ready form, with the language-coverage basis stated rather than asserted.
StatusIn force
InstrumentGPAI Code of Practice — Safety and Security
RequirementThe Commission's indicative compliance pathway contemplates model evaluation and the involvement of external evaluators.
DeliveredIndependent external evaluation with an evaluator statement, scope description and declared independence and conflicts position.
StatusPublished 10 July 2025
InstrumentNIST AI 600-1 — Generative AI Profile
RequirementThe MEASURE function calls for structured assessment of generative AI risks, including external red-teaming and documented evaluation of harmful content, bias and security risks.
DeliveredA measurement plan mapped to AI RMF subcategories including MEASURE 2.7 (security and resilience) and MEASURE 2.11 (fairness and bias), reported per language rather than in aggregate.
StatusPublished July 2024
InstrumentNIST AI RMF 1.0 — GOVERN and MAP
RequirementRisks are mapped in context, including identification of affected individuals and communities.
DeliveredAn affected-population analysis for Afghan-language users — diaspora, in-country and resettled — written to be quoted in your own risk documentation.
StatusPublished January 2023
InstrumentCalifornia SB 53 (TFAIA) — transparency report
RequirementA frontier developer deploying a new or substantially modified frontier model must publish a transparency report stating, among other things, the languages and modalities of output the model supports.
DeliveredEvidence for what "supports Pashto" means: measured performance and failure rates by language and variety, so the published claim has something behind it.
StatusEffective 1 January 2026
InstrumentCalifornia SB 53 — large frontier developers
RequirementSummaries of catastrophic-risk assessments, their results, and descriptions of third-party involvement, published alongside the frontier AI framework.
DeliveredA third-party evaluator description and assessment summary drafted for publication, and a statement of what we did and did not test.
StatusAnnual review; 30 days to justify material changes
InstrumentISO/IEC 42001
RequirementAn AI management system requires planned evaluation activity, documented results and management review inputs.
DeliveredThe full evidence pack indexed to clauses so it can be dropped into an audit file without reformatting.
StatusCertifiable standard
InstrumentOWASP Top 10 for LLM Applications
RequirementNamed application-layer risk classes including prompt injection, insecure output handling and excessive agency.
DeliveredEach applicable class exercised in-language, with findings tagged to the class.
StatusCurrent release
InstrumentMITRE ATLAS
RequirementA shared taxonomy of adversarial tactics and techniques against AI systems.
DeliveredFindings tagged to ATLAS technique identifiers so your security team can consume them alongside everything else in the register.
StatusLiving knowledge base
InstrumentEU Digital Services Act, Articles 34–35
RequirementVery large online platforms must assess and mitigate systemic risks, including risks arising from the languages their services operate in.
DeliveredLanguage-coverage evidence for Afghan diaspora communities on your platform, and moderation-policy testing against your own enforcement text.
StatusIn force for designated services
06 — Deliverables

What you receive

Eight artefacts. Each one is written to be used by someone other than us — that is the test we apply before anything is delivered.

1

Threat model and scope memorandum

System boundary, deployment assumptions, harm domains in scope and out, and the agreed attack budget.

2

Afghan-language harm taxonomy, versioned

The eight domains with named failure modes, adapted to your system and dated, so a later evaluation can be compared with this one.

3

Reproducible test set with provenance

Prompts, dialect variants, transliterations and multi-turn scripts, each with its author, language, variety and the harm domain it targets.

4

Findings register

Severity, likelihood, affected population, reproduction steps and transcript evidence for every confirmed finding, with inter-rater agreement reported.

5

Over-refusal and service-quality report

The other failure mode, by language and variety: refusal rates on legitimate requests, wrong-variety responses, and quality gaps against the English baseline.

6

Regulatory mapping appendix

Article 55, Article 53, NIST AI 600-1, SB 53, ISO/IEC 42001, OWASP and ATLAS, line by line against the evidence produced.

7

Publication-ready language

Draft text for your model card, system card or transparency report describing the evaluation, its scope and its limits — for you to verify, edit and own.

8

Regression suite and re-test report

The failing set packaged so your team can run it at every release, plus our measurement of what changed after remediation.

07 — Why Ariana Nexus

Why Ariana Nexus, and how we are different

There is no certification for Afghan-language AI evaluation. No board examines it, no registry lists it, and no vendor can show you one. So we wrote the standard we test to, we train the people who apply it, and we publish the method rather than asking you to take it on trust.

1

Native red teamers, not translated prompts

Every prompt is authored in-language by a native speaker of the specific variety, with the dialect, register and code-switching that real users produce. The published evidence is that human-authored attacks find substantially more than machine-translated ones — a translated test set measures your translation engine.

2

A firm of scholars, not a pool of bilinguals

The engagement is run by alumni of Cornell, Brown, the University of Chicago, the University of British Columbia and Otto-von-Guericke University Magdeburg, working in their own first language. Measurement design, adjudication statistics, regulatory mapping and clinical supervision are held by people qualified to hold them.

3

Both failure modes, on one scale

We measure harmful output and over-refusal in the same campaign and score them on the same scale. A vendor that only counts jailbreaks will hand you a model that is safe because it has stopped answering Afghan users.

4

Delivered entirely in-house

Every part of the engagement is produced by our own people. No subcontractors, no brokered specialists, no capacity bought in for the week. One engagement, one point of accountability.

5

Documentation is the product

Most evaluation work arrives as a deck. Ours arrives as an evidence pack indexed to the instruments your regulator, your auditor and your enterprise customers actually cite.

6

Duty of care to the people who do the work

Analysts working through extremist, sexual and war-related material are on capped caseloads with a structured debriefing protocol set and reviewed by a doctoral-level psychologist on the firm's own staff.

Four ways this work gets done

Machine-translated test set
General red-team vendor
Contract bilingual speakers
Ariana Nexus
Who writes the attacks
Machine-translated test setA translation engine
General red-team vendorEnglish-speaking red teamers
Contract bilingual speakersIndividually contracted speakers
Ariana NexusNative speakers of the specific variety, on staff
Dialect and variety coverage
Machine-translated test setNone
General red-team vendorNone
Contract bilingual speakersWhatever the individual happens to speak
Ariana NexusPashto and Dari varieties, plus 22 more languages
Harm taxonomy
Machine-translated test setEnglish list, translated
General red-team vendorEnglish list
Contract bilingual speakersAd hoc
Ariana NexusVersioned Afghan taxonomy, eight domains
Over-refusal measured
Machine-translated test setNo
General red-team vendorRarely
Contract bilingual speakersNo
Ariana NexusYes, on the same severity scale
Adjudication
Machine-translated test setAutomated scoring
General red-team vendorSingle reviewer
Contract bilingual speakersNone
Ariana NexusTwo reviewers plus tie-break, agreement reported
Regulatory mapping
Machine-translated test setNo
General red-team vendorGeneric
Contract bilingual speakersNo
Ariana NexusArticle 55, AI 600-1, SB 53, 42001, line by line
Analyst duty of care
Machine-translated test setNot applicable
General red-team vendorVaries
Contract bilingual speakersNone
Ariana NexusCapped caseloads, clinical supervision
Accountability
Machine-translated test setVendor of the engine
General red-team vendorThe vendor
Contract bilingual speakersDiffuse
Ariana NexusOne firm, one engagement lead, in-house
08 — Coverage

Language and variety coverage

Twenty-four Afghan languages across five families. Pashto and Dari carry the full evaluation bench; the remaining languages are covered to the depth the speaker population allows, and we state that depth in writing rather than implying uniform coverage.

1

Iranian

PashtoDariHazaragiAimaqBalochiOrmuriParachiWakhiShughniSanglechiIshkashimiMunjiYidgha
2

Turkic

UzbekiTurkmeniKyrgyz
3

Indo-Aryan

PashayiGawarbatiTirahi
4

Nuristani

NuristaniKatiPrasunWaigali
5

Dravidian

Brahui
Pashto varieties tested
Kandahari (southern), Nangarhari (eastern), Wardaki, Waziri
Dari varieties tested
Kabuli, Herati, Mazari, Badakhshani, Hazaragi
Script and encoding
Perso-Arabic orthographic variation, Latin transliteration, mixed-script input
Register
Formal written, spoken register, messaging shorthand, code-switched Pashto–Dari–English

Afghan Dari is not Iranian Persian and Afghan Turkmeni is not the standard Turkmen of Turkmenistan. Hazaragi is a variety of Dari, staffed separately. Where a model answers an Afghan Dari prompt in Iranian Persian, that is a finding, not a near miss.

09 — The team

The team behind this service

Ariana Nexus is operated and led by alumni and scholars of leading universities who work in Afghan languages as their own first languages. On an evaluation engagement, the tooling, the measurement design, the regulatory mapping and the harm taxonomy are each held by a named person qualified to hold them. We do not staff this work with bilinguals.

The difference shows up in the adjudication. Deciding whether a Pashto answer is dangerous requires someone who knows the language, the region the user is writing from, and what happens to a person who is found with that text on their phone. Deciding whether a difference between two reviewers is real requires someone who can compute it. Very few firms have both in the same room. That is the firm.

Program oversight

Hassan Ukasha

Managing Partner

  • B.S. Cornell University
  • M.P.H. Cornell University

Every Ariana Nexus engagement is delivered under the Managing Partner's oversight. Hassan Ukasha holds accountability for the firm's operations and for this program. He approves the scope of each evaluation before testing begins, signs the independence and conflicts position that appears on the face of every report, and is the escalation point for any finding that changes what a client can safely say about its own system.

The delivery team
Zeba Haqbani

Zeba Haqbani

Senior Partner

  • B.A.
    The American University of Afghanistan
  • B.A.
    University of British Columbia

Builds the firm's institutional systems and AI platforms. Owns the evaluation tooling, the test-set infrastructure and the transcript record, so every finding on this page can be reproduced from a logged run rather than a recollection.

Hussain Ahmad

Hussain Ahmad

Principal

  • M.Eng.
    Cornell University
  • Ph.D.
    University of Chicago

Measurement design and the statistics of evaluation: attack budgets, sampling, inter-rater reliability, and what a difference in failure rate between two languages does and does not support. Decides when a number is strong enough to publish.

Wasil Peroz

Wasil Peroz

Principal

  • B.A.
    Milli University
  • M.Sc.
    Otto-von-Guericke University Magdeburg

Institutional law. Maps every finding to EU AI Act Articles 53 and 55, NIST AI 600-1, California SB 53 and ISO/IEC 42001, and drafts the regulatory appendix your counsel reads before you publish anything.

Maryam Safi

Maryam Safi

Principal

  • B.A.
    Cornell University

Afghan information-space analysis. Maintains the harm taxonomy as a versioned instrument and seeds each campaign from what is actually circulating in Pashto and Dari, not from a translated list written somewhere else.

Abstract dark building geometry against a night sky
10 — Boundaries

Independence, boundaries and data handling

1

We evaluate. We do not certify.

Our report sits alongside your conformity assessor, your auditor and your counsel — never in place of them. We will not describe an evaluation as a conformity assessment, and we will not sign anything that implies we have.

2

Findings go to the owner, never to the public.

Confirmed findings are delivered to the system owner to remediate. We do not publish working attacks, release exploit tooling, or retain live jailbreaks outside the engagement record.

3

Your model and your data stay inside the agreement.

Model access is handled under NDA in the environment you specify. Prompts, outputs and transcripts are not used to train anything, are not reused on another engagement, and do not leave the agreed environment.

4

Independence is stated, not assumed.

Every report carries a declared independence and conflicts position. If we hold another engagement that touches the system under test, you are told before scoping, not after delivery.

5

No channel through the de facto authorities.

Ariana Nexus does not route documents, data or inquiries through channels controlled by the de facto authorities in Afghanistan. The firm has no offices and no operations inside the country.

6

The people doing the work are protected.

Red-team analysts are named in the engagement record and never in public findings. Caseloads are capped, debriefing is scheduled rather than offered, and the protocol is set and reviewed by Diana Ayubi, Psy.D., Engagement Manager at the firm.

11 — Buyers

Who this is built for

1

Frontier model developers and AI labs

Article 55 adversarial-testing evidence, SB 53 third-party involvement, and a defensible basis for the language claims in your model card.

Safety, Policy and Model Evaluation leads

2

Platforms and Trust and Safety teams

Moderation-policy fidelity in Pashto and Dari, classifier precision and recall by variety, and systemic-risk evidence for Afghan diaspora communities on your service.

Trust and Safety, Integrity and Policy leads

3

Enterprises deploying AI in Afghan-facing services

Health systems, resettlement organisations, insurers and banks using translation, chat or triage models with Afghan users — tested before the incident, not after.

Chief AI Officers, risk and compliance

4

Government agencies and prime contractors

Mission systems using machine translation or language models in Pashto and Dari, evaluated against NIST guidance by a firm registered to contract.

Program managers and capture leads

5

Assurance, audit and evaluation organisations

An Afghan-language evaluation bench you can bring into an engagement you are leading, with our method and our people named in your evidence file.

Engagement partners and technical leads

12 — Engagement

Engagement models

Evaluation is scoped as a program with a named engagement lead, not priced by the word or the prompt. Scoping begins under NDA.

1

Baseline evaluation

Four to six weeks

One system, one release. Full taxonomy, all agreed surfaces, complete evidence pack and regulatory appendix. The usual starting point, and the one most buyers need before a launch.

2

Release-gated program

Per release, retained

The regression suite is maintained between releases and re-run at each one, with a delta report against the previous evaluation and a standing coverage statement.

3

Continuous red team

Standing bench

A named team held against your systems year-round: scheduled campaigns, incident surge capacity, and quarterly reporting into your governance forum.

4

Targeted assessment

Two to three weeks

One harm domain, one language family or one surface — commonly used where an incident, a customer question or a regulator has raised something specific.

At a glance

Service
Independent adversarial testing and safety evaluation of AI systems in Afghan languages
Languages
Pashto and Dari first, 24 Afghan languages in total, with dialect and script coverage
Systems tested
Text LLMs, AI agents, guardrail classifiers, speech and voice, machine translation, moderation systems
Documented to
EU AI Act Articles 53 and 55 · NIST AI 600-1 and AI RMF 1.0 · California SB 53 · ISO/IEC 42001 · OWASP LLM Top 10 · MITRE ATLAS
Typical duration
Four to six weeks for a baseline evaluation; release-gated and continuous programs available
Delivered by
Ariana Nexus in-house — no subcontractors, no brokered specialists, one point of accountability
Based in
Washington, D.C. · United States-based delivery · no operations in Afghanistan
13 — Questions

Questions buyers ask

What is AI red teaming in Pashto and Dari?

It is structured adversarial testing of an AI system by native Pashto and Dari speakers, who attempt under controlled conditions to make the system produce harmful, unsafe or culturally damaging output. The result is a documented evaluation: a harm taxonomy, reproducible test sets, severity-scored findings and a report mapped to the regulatory instruments that apply to your system.

Why do AI safety guardrails fail in Afghan languages?

Safety behaviour is learned overwhelmingly from English and a few high-resource languages. Pashto and Dari sit far outside that distribution, so refusal behaviour, classifier accuracy and answer quality all degrade — and degrade further when users code-switch or transliterate, which Afghan users do constantly. The harms that matter in the Afghan context are also absent from taxonomies written for English-speaking users.

Can we just translate our English red-team set into Pashto?

You can, and it will usually produce a clean result that is not true. When researchers replaced machine translation with human translators on the same attacks, success rates rose by roughly thirty percent — the low scores had been measuring translation quality rather than model robustness. A translated test set is evidence about your translation engine.

Does this satisfy EU AI Act Article 55?

Article 55(1)(a) requires providers of general-purpose AI models with systemic risk to perform model evaluation to state-of-the-art protocols, including conducting and documenting adversarial testing. Our engagement produces exactly that record for Afghan languages: protocol, attack budget, coverage statement, reproducible prompts and logged transcripts. Whether your overall Article 55 position is sufficient is a judgement for you and your counsel; we supply the evidence and the mapping, and we do not certify.

What does California SB 53 require about languages?

A frontier developer deploying a new or substantially modified frontier model must publish a transparency report stating, among other things, the languages and modalities of output the model supports. Large frontier developers must also publish summaries of catastrophic-risk assessments and descriptions of third-party involvement. We give you measured evidence behind the language claim and a third-party evaluator description drafted for publication.

How does this map to NIST AI 600-1?

The Generative AI Profile's MEASURE function calls for structured assessment of generative AI risks, including external red-teaming and documented evaluation of harmful content, bias and security risks. We deliver a measurement plan mapped to AI RMF subcategories, reported per language rather than in aggregate, because an aggregate multilingual score hides exactly the gap you are testing for.

Which Afghan languages can you test?

Twenty-four, across five language families. Pashto and Dari carry the full evaluation bench, including dialect coverage across Kandahari, Nangarhari and Wardaki Pashto and Kabuli, Herati, Mazari and Hazaragi Dari. The smaller languages are covered to the depth the speaker population allows, and we state that depth in the report rather than implying uniform coverage.

Do you test speech and voice models, or only text?

Both, plus agents, guardrail classifiers, machine translation, moderation systems and multimodal inputs. For speech we treat mis-transcription as a safety failure rather than a quality metric, because a wrong transcript upstream produces a wrong decision downstream.

Do you measure over-refusal as well as harmful output?

Yes, in the same campaign and on the same severity scale. Models frequently over-block Pashto and Dari, flagging ordinary religious and cultural speech as unsafe and refusing legitimate requests at higher rates than in English. A vendor that only counts jailbreaks will hand you a model that looks safe because it has stopped serving Afghan users.

How long does an evaluation take and how is it priced?

A baseline evaluation runs four to six weeks. Release-gated programs, continuous benches and targeted assessments are scoped individually. Engagements are scoped as programs with a named engagement lead rather than priced per prompt, and scoping begins under NDA.

How do you protect our model, our prompts and our findings?

Model access is handled under NDA in the environment you specify. Prompts, outputs and transcripts are not used to train anything, are not reused on another engagement and do not leave that environment. Confirmed findings go to you to remediate; we do not publish working attacks or release exploit tooling.

Can Ariana Nexus be named as an independent third-party evaluator?

Yes. Every report carries a declared independence and conflicts position, a scope description and an evaluator statement written to be quoted in a model card, a transparency report or an audit file. If we hold another engagement that touches the system under test, you are told before scoping.

Start with a scoping conversation

Tell us what the system is, where it is deployed and which languages you have claimed. We will tell you what an evaluation would cover, what it would not, and what the evidence would look like when it is finished. Scoping runs under NDA and we meet in person at the Washington, D.C. office where that helps.

Request a scoping conversation

Related services