Skip to main content

AI Governance Documentation for Pashto, Dari and 22 More Afghan Languages

EU AI Act, California SB 53 and ISO/IEC 42001 evidence

Ariana Nexus prepares the Afghan-language evidence that AI governance files are usually missing: what your model or system was tested on in Pashto and Dari, how it performed, where it fails, who reviewed it and how it is monitored. Each document is written to the EU AI Act, California SB 53 or ISO/IEC 42001 requirement it supports.

See what the evidence file contains
Fig. 1 · One dossier, three submissions. Select a module to read its documents. The same modules are lifted into an EU technical file, a California transparency report and an ISO/IEC 42001 management system.

What is AI governance documentation for language coverage?

AI governance documentation for language coverage is the written evidence of how an AI model or system was built, tested, limited and monitored in a specific language. For Afghan languages it records the Pashto and Dari training and test data, evaluation results by language and dialect, known failure modes, human oversight arrangements and incident history, in the structure that the EU AI Act, California SB 53 and ISO/IEC 42001 ask for.

Frameworks
EU AI Act (Regulation (EU) 2024/1689, as amended by Regulation (EU) 2026/1744), California SB 53, ISO/IEC 42001:2023. Also mapped on request: NIST AI RMF and AI 600-1, the EU General-Purpose AI Code of Practice, ISO/IEC 42005, Digital Services Act Articles 34 and 35
Languages
Pashto and Dari first, then 22 more Afghan languages by published readiness tier
Built for
General-purpose AI model providers and frontier developers, providers and deployers of high-risk AI systems, online platforms, organizations certifying to ISO/IEC 42001, and their auditors, notified bodies and counsel
Who does the work
University-educated native speakers tested in the variety they review, led by alumni and scholars of Cornell University, the University of Chicago, the University of British Columbia and Otto-von-Guericke University
What you receive
A controlled, versioned evidence file of ten documents, a requirement-to-evidence crosswalk and a machine-readable index
Delivery
In-house, from Washington, D.C. No subcontractors and no crowd platforms
Boundary
Evidence and documentation. Not legal advice, not certification, not a conformity assessment

How AI governance documentation for Pashto and Dari works

From the records you already hold to one evidence file that three frameworks can read.

Fig. 2
From the records you hold to one file that three frameworks can read
1 · What you hold
  • Model cards and system cards
  • Data sheets
  • Evaluation logs
  • Risk registers
  • Public language claims
2 · What we add
  • Native-written Pashto and Dari tests
  • Data review by variety and script
  • Safety evaluation against English
  • Drafting to the requirement
  • Independent second review
3 · The evidence file
  • Ten controlled documents
  • Requirement-to-evidence crosswalk
  • Evaluator record
  • JSON and CSV index
  • Version history
4 · Who reads it
  • EU AI Act: AI Office, notified body, deployer
  • California SB 53: the public, the Attorney General
  • ISO/IEC 42001: certification body
  • Your own counsel and auditors
Evidence and documentation only. Not legal advice, not certification, not a conformity assessment.

Most AI governance files are silent on Pashto and Dari

Conformance is not coverage.

A model card says the model supports more than a hundred languages. The technical documentation reports accuracy on English benchmarks. The risk assessment lists multilingual performance as a single line. None of that tells a regulator, an auditor or a deployer what happens when a person writes to the system in Pashto or Dari.

That gap is now a documentation problem, not only a quality problem. California SB 53 requires frontier developers to publish the languages a model supports. The European Commission’s training-content template asks general-purpose AI providers to describe the linguistic characteristics of their data. The EU AI Act requires high-risk systems to be trained and tested on data that represents the people and settings they are used in, and to declare accuracy and known limitations. ISO/IEC 42001 asks for an assessment of the system’s impact on individuals and groups.

A management system can conform to a standard and still hold no evidence about an entire language community. Conformance is not coverage. This service supplies the coverage.

Fig. 3
What governance files usually say about languages, and what an assessor needs to see
What the file usually says
What an assessor needs to see
“Supports 100+ languages.”
Which Afghan languages, to what standard, tested how, by whom
Accuracy on English or machine-translated benchmarks
Results on native-written Pashto and Dari test items, by variety and script
“Multilingual data” in the training summary
Sources, volume, provenance and variety balance of the Pashto and Dari data; Dari separated from Iranian Persian
One risk line: “performance may vary by language”
Named risks for Afghan-language users, with severity, controls and residual risk
Human oversight “by trained staff”
Who can read Pashto and Dari output, how they were tested, when the system hands off to them
Incident process written for English
How an Afghan-language failure is detected, classified, reported and closed
Ten documents in five modules

What we deliver: the Afghan-language evidence file

Ten controlled documents. Each one can be lifted into your technical documentation, framework, transparency report or management system without rewriting.

M1: Regional requirements

One cover sheet per framework, and the crosswalk that ties every page of evidence to it.
EU AI Act · California SB 53 · ISO/IEC 42001
D-10

Requirement-to-evidence crosswalk and evaluator record

Every document, table and test traced to the article, section, clause or control it supports, with a record of who did the work, their qualifications, their independence and any conflict.
EU AI Act
Art. 11, Annex IV; Art. 17; Art. 18
SB 53
§ 22757.12(b): annual review
ISO/IEC 42001
Clauses 7.5, 9.2, 9.3; Statement of Applicability

Where each module is filed

Submission
Status
Next date
EU AI Act
In force. General-purpose AI obligations apply now. High-risk obligations apply from 2 December 2027 (Annex III) and 2 August 2028 (Annex I products).
2 December 2027
California SB 53
In effect since 1 January 2026. Civil penalties of up to $1 million per violation, enforced by the Attorney General.
Each new or substantially modified frontier model
ISO/IEC 42001
Published December 2023. Voluntary, and certifiable by accredited certification bodies. Not a harmonised standard under the EU AI Act.
Your stage 1 and stage 2 audit dates

M2: Summaries and public statements

What you say in public and to deployers: coverage, limits and safety summaries.
D-01

Language coverage statement

Each Afghan language and variety marked supported, limited or not supported, with the test basis for every claim and the wording to use in public disclosures.
EU AI Act
Art. 13(3)(b); Art. 53(1)(a)–(b), Annexes XI and XII
SB 53
§ 22757.12(c): languages supported
ISO/IEC 42001
A.8.2, A.9.4
D-06

Safety evaluation summary

Refusal consistency, jailbreak resistance and harmful-content handling in Pashto and Dari compared with English, in the summary form that frameworks and transparency reports require.
EU AI Act
Art. 55(1)(a); Annex XI, Section 2
SB 53
§ 22757.12(c): risk assessment summaries and third-party evaluators
ISO/IEC 42001
A.6.2.4; NIST AI 600-1
D-08

Instructions for use and transparency text

Plain statements of performance and limitation by language for deployers, downstream providers and the public: instructions for use, model card and system card language sections, transparency report entries.
EU AI Act
Art. 13; Art. 53(1)(b), Annex XII
SB 53
§ 22757.12(c): intended uses and restrictions
ISO/IEC 42001
A.8.2, A.8.5

M3: Data

What the system learned from and was tested on in Pashto and Dari.
D-03

Data governance record

Pashto and Dari training, validation and test data: sources, provenance, collection and consent basis, balance by variety and script, known gaps, bias examination and preparation steps.
EU AI Act
Art. 10(2)–(4); Annex IV(2)(d); Art. 53(1)(d)
SB 53
Not addressed by this instrument
ISO/IEC 42001
A.7.2–A.7.6, A.4.3

M4: Testing

How it performs, by language, variety, script and task.
D-02

Evaluation and test report

Test design, native-written test sets, metrics, results by language, variety, script and task, inter-rater agreement and limitations.
EU AI Act
Art. 15; Annex IV(2)(g), (3), (4); Art. 55(1)(a)
SB 53
§ 22757.12(a): assessment results
ISO/IEC 42001
A.6.2.4; clause 9.1

M5: People and operations

Who is affected, who oversees, how failures are found and reported.
D-04

Language-specific risk register

Risks that arise because the user writes or speaks an Afghan language: mistranslation in a high-stakes setting, unsafe output that English filters would block, over-refusal, dialect and gender errors, misidentified language. Severity, likelihood, controls and residual risk.
EU AI Act
Art. 9; Art. 55(1)(b)
SB 53
§ 22757.12(a): risk assessment and mitigation
ISO/IEC 42001
Clauses 6.1.2, 6.1.3, 8.2, 8.3
D-05

Impact assessment input

Who is affected and how: Afghan asylum seekers, refugees, patients, students, workers and families; the consequence of a wrong output in each setting; safeguards. Drafted to drop into a fundamental rights impact assessment or an AI system impact assessment.
EU AI Act
Art. 27; Art. 10(2)(f)–(g)
SB 53
Not addressed by this instrument
ISO/IEC 42001
Clauses 6.1.4, 8.4; A.5.2–A.5.5; ISO/IEC 42005
D-07

Human oversight and escalation procedure

Who reviews Afghan-language output, the qualification they must hold, what they check, when the system must hand off to a person and how overrides are logged.
EU AI Act
Art. 14; Art. 26(2)
SB 53
§ 22757.12(a): internal governance
ISO/IEC 42001
A.3.2, A.4.6, A.9.2
D-09

Monitoring and incident plan

What to log, what to sample, the thresholds that trigger review, and how an Afghan-language failure is classified, reported and closed.
EU AI Act
Art. 72; Art. 73; Art. 55(1)(c)
SB 53
§ 22757.13: critical safety incidents
ISO/IEC 42001
Clause 9.1; A.6.2.6, A.6.2.8, A.8.4

Delivered as controlled Word and PDF documents with version history, plus a machine-readable index (JSON and CSV) that imports into your document control or GRC system.

How the evidence maps to the EU AI Act, SB 53 and ISO/IEC 42001

One body of evidence, three readers. The same Pashto and Dari test results answer a different question in each framework.

Fig. 4
Requirement-to-evidence crosswalk: one body of evidence, three readers
Evidence area
EU AI Act
California SB 53
ISO/IEC 42001
Languages supported and their limits
Art. 13(3)(b); Art. 53(1)(a)–(b), Annexes XI and XII
§ 22757.12(c): languages supported, intended uses, restrictions
A.8.2, A.9.4
Training and test data
Art. 10(2)–(4); Annex IV(2)(d); Art. 53(1)(d) training-content summary
Not addressed directly
A.7.2–A.7.6, A.4.3
Accuracy, robustness and testing
Art. 15; Annex IV(2)(g), (3), (4)
§ 22757.12(a): assessments and their results
A.6.2.4; clause 9.1
Risk assessment and mitigation
Art. 9; Art. 55(1)(a)–(b)
§ 22757.12(a), (c): catastrophic-risk assessment and summaries
Clauses 6.1.2, 6.1.3, 8.2, 8.3
Impact on people and groups
Art. 27; Art. 10(2)(f)–(g)
Not addressed directly
Clauses 6.1.4, 8.4; A.5.2–A.5.5
Human oversight
Art. 14; Art. 26(2)
§ 22757.12(a): internal governance practices
A.3.2, A.4.6, A.9.2
Monitoring and incidents
Art. 72; Art. 73; Art. 55(1)(c)
§ 22757.13: critical safety incident reports
Clause 9.1; A.6.2.6, A.6.2.8, A.8.4
Records and document control
Art. 11, 12, 18; Annex IV
§ 22757.12(b), (e): annual review; no false or misleading statements about risk
Clauses 7.5, 9.2, 9.3

References are a navigation aid prepared on 19 September 2026 against the current legal texts. They are not legal advice. Your counsel decides which obligations apply to you.

When the obligations apply

Checked against the Official Journal and the California code on 19 September 2026.

Fig. 5
The dates that govern the file, verified on 19 September 2026
  1. ISO/IEC 42001:2023 published. Certification is available now.
    In effect
  2. EU AI Act enters into force.
    In effect
  3. EU obligations for general-purpose AI model providers apply (Articles 53 and 55). Code of Practice and training-content template published in July 2025.
    In effect
  4. California SB 53 takes effect: frontier AI frameworks, transparency reports listing languages supported, incident reporting.
    In effect
  5. Regulation (EU) 2026/1744, the Digital Omnibus on AI, enters into force and moves the high-risk dates.
    In effect
  6. European Commission enforcement powers over general-purpose AI providers begin. Article 50 transparency obligations apply.
    In effect now
  7. General-purpose AI models placed on the market before 2 August 2025 must comply.
    Upcoming
  8. High-risk obligations apply to stand-alone Annex III systems: asylum and migration, public benefits, education, employment, justice.
    Upcoming
  9. High-risk obligations apply to AI in Annex I regulated products, including medical devices.
    Upcoming
Filled marker: in effect. Gold: in effect now. Open: upcoming. Each document records the date its references were last checked.

Why Pashto and Dari are hard to evidence

Both are low-resource languages. That changes what can be assumed from English results, and what has to be shown directly.

Fig. 6
How little Pashto there is on the web
English
40.45%
German
5.91%
Persian, including Afghan DariDari is not identified separately
0.76%
Urdu
0.032%
Uzbek
0.031%
Pashto
0.0043%
Turkmen
0.0033%
Share of web pages by primary language, Common Crawl CC-MAIN-2026-34 (August 2026), identified by CLD2. Bars are drawn to a linear scale: at this width the Pashto bar is less than a tenth of a pixel, which is why it cannot be seen.
Fig. 7
Safeguards fail most often in low-resource languages
English (no translation)
0.96%
High-resource languages
10.96%
Mid-resource languages
21.92%
Low-resource languages
79.04%
Share of harmful requests that bypassed GPT-4’s safeguards when translated, by language resource level. Yong, Menghini and Bach, “Low-Resource Languages Jailbreak GPT-4”, 2023. Pashto and Dari were not among the languages tested.
01

Very little web text

In Common Crawl’s August 2026 crawl, 40.45% of pages were identified as English and 0.0043% as Pashto: about 9,400 English pages for every Pashto page.
Common Crawl, CC-MAIN-2026-34, primary language by CLD2
02

Dari is not counted separately

Standard language identification has no Dari class. Afghan Dari is folded into Persian (0.76% of pages), a category dominated by Iranian sources. A data sheet that says “Persian” says nothing about the vocabulary, spelling and register used in Afghanistan.
Common Crawl, CC-MAIN-2026-34
03

Safety training does not transfer evenly

In a 2023 Brown University study, translating harmful requests into low-resource languages bypassed GPT-4’s safeguards 79% of the time, against under 1% for the same requests in English. The study did not test Pashto or Dari; it shows why safety claims need evidence in each language.
Yong, Menghini and Bach, 2023 (arXiv:2310.02446)
04

Machine-translated benchmarks mislead

A test set translated from English by machine carries the translation system’s own errors and none of the local content: no Afghan names, places, institutions, law or idiom. The score measures the translation as much as the model.
05

Several varieties, more than one way of writing

Pashto and Dari each have regional varieties, and Hazaragi is a variety of Dari with its own vocabulary. People write in Arabic script, in Latin letters, and mixed with English, German or Urdu. A result for formal written Pashto says little about a message typed in Latin letters by a teenager in Hamburg.
06

Few reviewers can read the output

Human oversight only works if the overseer understands the output. Most governance, safety and audit teams have no one who reads Pashto or Dari, and no test for the people they borrow.

Where Afghan-language coverage decides the outcome

More than forty years of war and displacement have placed Afghan communities in front of AI systems in asylum offices, hospitals, schools, job centres and social platforms across Europe, North America and Australia.

01

Asylum, migration and border systems

Afghans are consistently among the largest groups of asylum applicants in the European Union. AI used to examine applications or assess evidence is listed as high-risk in Annex III. A translation error in a Pashto interview transcript, or a Dari speaker’s language identified as Iranian Persian, can change a decision about a person’s protection.
02

Healthcare and patient communication

Hospitals and health apps use AI to translate discharge instructions, triage messages and answer symptom questions for Afghan patients. Pashto and Dari expressions for pain, distress and trauma do not map cleanly onto English clinical terms, and a system tuned on English misses them.
03

War, displacement and trauma content

Everyday conversation in Pashto and Dari carries the Afghan war, loss and exile. Safety systems trained on English either block ordinary talk about conflict and grief, or miss real signs of self-harm and threat because they are phrased in local idiom. Both failures belong in the risk file.
04

Public services, education and employment

Benefit eligibility, school placement, language testing and hiring tools are Annex III uses. The Afghan diaspora in Germany, the United States, Canada, the United Kingdom and Australia meets these systems in a second or third language, or through machine translation.
05

Online platforms and content moderation

Pashto and Dari moderation has to separate news, poetry, religious speech and political argument from incitement, extremist propaganda and harassment of women and minorities. Over-removal silences a community; under-removal leaves it exposed.
06

Frontier model safety

A refusal that holds in English and fails in Pashto is not a mitigation. If a transparency report lists a language as supported, the safety claims in the same report should hold in that language.

Who needs Afghan-language governance evidence

01

General-purpose AI model providers and frontier developers

You publish a list of supported languages, a training-content summary and a transparency report. You need evidence behind the Pashto and Dari lines, or a documented basis for marking them limited or unsupported.
EU AI Act Articles 53 and 55 · SB 53 § 22757.12
02

Providers of high-risk AI systems

Your system is used with Afghan-language speakers in asylum and migration, public benefits, healthcare, education, employment or justice. You need data governance, accuracy and limitation statements that cover them.
EU AI Act Articles 9 to 15 and Annex IV
03

Deployers: public authorities, hospitals, agencies and NGOs

You are putting an AI translation, triage, screening or case-management tool in front of Afghan clients. You need a fundamental rights impact assessment and an oversight procedure that account for language.
EU AI Act Articles 26 and 27
04

Online platforms and trust and safety teams

Your systemic risk assessment has to take regional and linguistic aspects into account. You need documented evidence of how moderation systems perform on Pashto and Dari content.
Digital Services Act Articles 34 and 35
05

Organizations certifying to ISO/IEC 42001

Your AI management system has multilingual systems in scope. You need impact assessment, data quality and validation records that a certification body can sample.
ISO/IEC 42001 clauses 6.1.4 and 8.4, Annex A.5, A.6, A.7
06

Auditors, notified bodies, certification bodies and counsel

You are assessing a file that contains Afghan-language material you cannot read. You need an independent specialist to examine it and report in terms you can rely on.
Expert input to an audit, assessment or legal review

How we produce the evidence

Six steps, each with a written output you can inspect before the next begins.

  1. Step 1

    Scope and regulatory map

    Confirm your role (provider, deployer, frontier developer, certification candidate), the systems in scope, the languages you claim, and the exact articles, sections and clauses the evidence must serve.
    Output: Signed scope and crosswalk skeleton · Client and Ariana Nexus
  2. Step 2

    Evidence inventory and gap analysis

    Read what you already hold (model cards, data sheets, evaluation logs, risk registers) and mark what is missing, untested or unreadable for Afghan languages.
    Output: Gap register · Ariana Nexus
  3. Step 3

    Native evaluation and data review

    Build or select test sets written by native speakers, run the evaluation, and review data samples for variety, script, quality and provenance. Every item is scored twice; disagreements are adjudicated.
    Output: Results with agreement statistics · Ariana Nexus
  4. Step 4

    Drafting to the requirement

    Write each document in the structure its reader expects: Annex IV sections, framework and transparency report entries, ISO documented information. Numbered claims, every figure traceable to its source.
    Output: Draft evidence file · Ariana Nexus
  5. Step 5

    Independent second review

    A reviewer who did not run the work checks each claim against its source, each legal reference against the current text, and the language content against the firm’s Cultural Compliance Bureau standards.
    Output: Review record and sign-off · Second reviewer
  6. Step 6

    Handover and update cycle

    Deliver into your document control or GRC system with the machine-readable index. Re-test and re-issue on each model release, material change or annual review.
    Output: Issued file, version 1.0 · Client and Ariana Nexus
Fig. 8
Who does what at each step, and the written output
Lane
1Scope and regulatory map
2Evidence inventory and gap analysis
3Native evaluation and data review
4Drafting to the requirement
5Independent second review
6Handover and update cycle
You
Confirm your role, systems and the languages you claim
No action at this step
No action at this step
No action at this step
No action at this step
Receive the file into document control
Ariana Nexus
Signed scope and crosswalk skeleton
Gap register
Results with agreement statistics
Draft evidence file
No action at this step
Issued file, version 1.0, with the index
Second reviewer
No action at this step
No action at this step
No action at this step
No action at this step
Review record and sign-off
No action at this step
Black: Ariana Nexus. Gold: the independent second reviewer. White: your team.

A first evidence file for Pashto and Dari on one system typically takes six to ten weeks from a signed scope. Updates for a new model version are faster, because the test sets, procedures and crosswalk already exist.

Ways to engage

Engagement
Duration
What you provide
What you receive
Gap assessment
Two to three weeks
Your existing file, model cards and public language claims.
A written gap register against the three frameworks and a plan to close it.
Full evidence file
Six to ten weeks
Model or system access, existing evaluation and data records.
The ten documents for one system, in Pashto and Dari, with crosswalk and index.
Retained evidence program
Annual
Release calendar and change notices.
Re-test on each release, incident support, annual review inputs, extension to further languages.
Expert to an auditor or counsel
By matter
The Afghan-language material in the file you are assessing.
An independent written report in terms a non-reader of the language can rely on.

Every engagement is scoped and priced as a whole, with a named partner accountable for it. We do not take per-item review work or place unnamed reviewers in another firm’s pipeline.

Why Ariana Nexus

Governance consultancies know the frameworks and cannot read the languages. Language vendors can read the languages and do not know what an auditor will sample. This firm was built to do both.

What we bring

01

Afghan languages are the entire practice

Twenty-four languages across five families. Not one line in a list of two hundred.
02

Scholars, not a crowd

Work is led by alumni and scholars of Cornell University, the University of Chicago, the University of British Columbia and Otto-von-Guericke University. Every evaluator holds a university degree and is tested in the specific variety they review.
03

The standard where no certification exists

No U.S. certification exists for Pashto or Dari interpreters, translators or AI evaluators, so there is no credential to check. We set the standard and train to it, for our own specialists and for outside interpreters. Each reviewer’s result goes into the evaluator record, and we assess deployer staff who oversee Pashto or Dari output under Article 26(2).

How we deliver

04

Written for the person who audits it

Numbered claims, traceable figures, legal references checked against the current text at every issue.
05

Produced in-house

Every part is produced by our own people. One engagement, one point of accountability. No subcontractors, no crowd platforms.
06

A clean data path

Non-disclosure agreement before any data. Nothing is routed through channels controlled by the de facto authorities in Afghanistan. GDPR and UK GDPR terms for European clients.

How we are different

07

Honest coverage

Where a language or variety cannot be evidenced, the file says “not supported” rather than implying it.
08

Alongside your certifier and counsel

Never in place of them. We supply the evidence they cannot produce and do not claim the judgments that are theirs.
Criterion
General AI governance consultancy
Language vendor or crowd platform
Ariana Nexus
Reads and evaluates Pashto and Dari output
No. Relies on the client’s own claims
Yes, through unnamed reviewers of unknown qualification
Yes. University-educated native specialists, tested in the variety
Knows what the framework asks for
Yes
No
Yes. Every document mapped to article, section, clause and control
Test material
English or machine-translated
Translated to order
Written natively for the Afghan context
What an auditor can sample
English evidence only
A spreadsheet of labels
A versioned file with traceability, agreement statistics and reviewer record
Accountability
Partner-led
Project manager and an anonymous crowd
Partner-led, in-house, one point of accountability

We do not build or sell AI models, governance software or certification. We do sell evaluation, red teaming and data services; where an evidence file reports on work we performed, the file says so.

What this service does not do

  • It does not give legal advice or an opinion on whether you comply.
  • It does not certify anything. ISO/IEC 42001 certificates come from accredited certification bodies; EU conformity assessment belongs to the provider or a notified body.
  • It does not issue a badge, seal or “compliant” label. No accredited scheme exists for language coverage, so we do not invent one.
  • It does not write your English-language governance program. It supplies the Afghan-language evidence that belongs inside it.
  • It does not adjust findings. If the model fails in Pashto, the file says so, and says where.
پښتودری

Afghan languages covered

All 24 appear in the coverage statement. We publish readiness tiers rather than a flat claim, because speaker availability differs by two orders of magnitude across this list.

Fig. 9
Readiness tiers: what the evidence file can contain today, by language
Tier 1
Pashto, Dari (including the Hazaragi variety)
All ten documents
The full ten-document evidence file, with native evaluation, data review and safety testing
Tier 2
Uzbeki, Turkmeni, Balochi, Pashayi, and the Nuristani group (Nuristani, Kati, Prasun, Waigali)
Three documents, on an agreed scope
Coverage statement, data review and evaluation on an agreed scope, with panel size stated in the file
Tier 3
Aimaq, Ormuri, Parachi, Tirahi, Gawarbati, Kyrgyz, Brahui, and the Pamir group (Wakhi, Shughni, Sanglechi, Ishkashimi, Munji, Yidgha)
Coverage statement and expert review
Coverage statement and expert review scoped to verified speaker availability. Where nothing can be evidenced, the file records “not evidenced”
Iranian · 13
Pashto, Dari, Hazaragi (a variety of Dari), Aimaq, Balochi, Ormuri, Parachi, Wakhi, Shughni, Sanglechi, Ishkashimi, Munji, Yidgha
Turkic · 3
Uzbeki, Turkmeni, Kyrgyz
Indo-Aryan · 3
Pashayi, Gawarbati, Tirahi
Nuristani · 4
Nuristani (Ashkun group), Kati, Prasun, Waigali
Dravidian · 1
Brahui

Dari deliverables are Afghanistan Dari, not Iranian Persian. Turkmeni deliverables are Afghan Turkmeni, not the standard Turkmen of Turkmenistan. Hazaragi is a variety of Dari and is staffed separately because the lexical distance is large enough to change a result.

The team behind the evidence

Ariana Nexus is led and operated by alumni and scholars of leading universities who are native speakers of the languages they assess. The people who read the Pashto and Dari output are the people who write the findings.

Hassan Ukasha, Managing Partner, Ariana Nexus
Program oversight

Hassan Ukasha

Managing Partner, Ariana Nexus, Washington, D.C.
B.S.
Cornell University
M.P.H.
Cornell University

Hassan Ukasha oversees the firm’s operations and its AI governance documentation program. He is the accountable partner on every evidence file Ariana Nexus issues: he approves the scope, signs the independence and conflicts declaration, and is the point of escalation for your general counsel, notified body or certification auditor.

B.Sc.
University of British Columbia
Zeba Haqbani

Zeba Haqbani

Senior Partner
M.Eng.
Cornell University
Ph.D.
University of Chicago
Hussain Ahmad

Hussain Ahmad

Principal
B.A.
Milli University
M.Sc.
Otto-von-Guericke University Magdeburg
Wasil Peroz

Wasil Peroz

Principal
B.A.
Cornell University
Maryam Safi

Maryam Safi

Principal

Who is accountable for each part of your file

Part of the evidence file
Accountable
What they sign off
Scope, independence and issue
Hassan Ukasha
The scope, the independence and conflicts declaration, and every issued version of the file.
Document control and the machine-readable index
Zeba Haqbani
The version history, the JSON and CSV index and the data-handling record.
Evaluation design and statistics
Hussain Ahmad
The sampling plan, the agreement statistics and every reported figure.
Requirement mapping and legal references
Wasil Peroz
The crosswalk, with every reference checked against the current legal text.
Public-sector and high-risk deployments
Maryam Safi
The evidence for asylum, migration and public-service systems.
Pashto and Dari scoring and review
Native specialists, named in the evaluator record
Every scored item. A second reviewer adjudicates every disagreement.

How this team works, and why it is different

Read, not relayed

Nothing passes through a project manager who cannot read the language. The specialist who evaluates the output drafts the finding.

Tested in the variety

A reviewer of Kandahari Pashto is assessed in Kandahari Pashto. A degree and a native language are entry conditions, not the qualification.

Two disciplines on every file

Each file is signed by a language lead and by a methods or regulatory lead, so that both the Pashto and the statistics are right.

Named and accountable

Every report names the partners accountable for it and records each evaluator’s qualifications under an evaluator ID. Contributors are pseudonymised for their safety; their credentials are not hidden.
A single Afghan flag standing among rows of United States flags on the lawn at Cornell University.

Questions buyers ask

Does the EU AI Act require testing in every language a system supports?

Not in those words. The Act does not list languages. It requires high-risk systems to use training, validation and test data that is relevant and sufficiently representative for the intended purpose and the setting of use (Article 10), to declare accuracy levels and known limitations (Articles 13 and 15), and to document testing in the Annex IV technical file. If Afghan-language speakers are within the intended purpose, the file needs evidence that covers them, or a stated limitation.

What does California SB 53 require about languages?

Before deploying a new or substantially modified frontier model, a frontier developer must publish a transparency report that includes the languages the model supports, its intended uses and its restrictions. Large frontier developers must also summarise their catastrophic-risk assessments and the role of third-party evaluators. SB 53 prohibits materially false or misleading statements about catastrophic risk and its management, which is one reason safety claims should be evidenced in every language the report lists. The law has applied since 1 January 2026.

How does Afghan-language evidence fit into ISO/IEC 42001 certification?

ISO/IEC 42001 certifies the management system, not the model. An auditor samples documented information: the AI system impact assessment (clause 6.1.4 and Annex A.5), data quality and provenance records (A.7), verification and validation (A.6.2.4) and information for users (A.8.2). Where a system in scope is used in Pashto or Dari, those records need language-specific content. We supply it; your certification body audits it.

Does ISO/IEC 42001 certification mean EU AI Act compliance?

No. ISO/IEC 42001 is a voluntary management system standard and is not a harmonised standard under the EU AI Act. It gives a governance structure that much of the Act’s documentation can sit in, but the Act’s specific obligations still have to be met and evidenced separately.

Are Pashto and Dari low-resource languages, and why does that matter for compliance?

Yes. In Common Crawl’s August 2026 crawl, 0.0043% of pages were identified as Pashto against 40.45% English, and Dari is not identified separately from Persian at all. Models see little Afghan-language data, safety training transfers unevenly and public benchmarks are scarce, so performance and safety claims cannot be assumed from English results. They have to be evidenced directly.

Can machine-translated benchmarks be used as evidence?

They can be disclosed, but they are weak evidence. A machine-translated test set carries the translation system’s errors and contains no Afghan names, places, institutions or idiom, so it measures translation quality as much as model quality. We use test items written by native speakers and state the method in the report so an assessor can judge its weight.

Do you provide legal advice or certify compliance?

No. Ariana Nexus is not a law firm, a notified body or a certification body. We produce Afghan-language evidence and documentation and map it to the requirements it supports. Your counsel decides what the law requires of you; your certification body or notified body decides conformity. We work alongside them, never in place of them.

Who does the work?

University-educated native speakers tested in the specific variety they review, led by partners and principals who are alumni and scholars of Cornell University, the University of Chicago, the University of British Columbia and Otto-von-Guericke University. No U.S. certification exists for Pashto or Dari, so every reviewer is assessed against the firm’s written standard and the result is recorded in the evaluator record. All work is done in-house. We do not use subcontractors or crowd platforms.

Which Afghan languages do you cover?

All 24 appear in the coverage statement, in three published readiness tiers. Pashto and Dari, including the Hazaragi variety, receive the full ten-document evidence file today. Uzbeki, Turkmeni, Balochi, Pashayi and the Nuristani group are evidenced on an agreed scope. The smallest languages, such as Ormuri, Parachi, Tirahi and the Pamir group, are scoped to verified speaker availability, and the file says “not evidenced” where that is the truth.

We deploy AI but do not build it. Do we need this?

Often, yes. Under the EU AI Act, deployers of high-risk systems must use them according to the instructions, assign competent human oversight and, for public bodies and certain services, complete a fundamental rights impact assessment before first use. If your clients, patients or applicants include Afghan-language speakers, those steps need language-specific content. We also review a vendor’s language claims before you buy.

How long does it take?

A gap assessment takes two to three weeks. A first full evidence file for Pashto and Dari on one system typically takes six to ten weeks from a signed scope, depending on how much evaluation already exists. Updates for a new model version are faster.

How are our model, data and documents protected?

Work begins under a mutual non-disclosure agreement. Access is limited to the named engagement team, data stays in the environment agreed in the scope, and contributors are pseudonymised in deliverables. Ariana Nexus routes no data, documents or inquiries through channels controlled by the de facto authorities in Afghanistan. European engagements run under GDPR and UK GDPR terms.

How is this different from your LLM evaluation and red teaming services?

Those services produce the measurements. This service turns measurements, ours or your own, into the governance record: coverage statements, data and risk files, oversight procedures and the crosswalk that auditors and regulators read. Many clients start here with a gap assessment, then commission the evaluation the gaps call for.

When do the EU AI Act obligations apply?

Obligations for general-purpose AI model providers have applied since 2 August 2025, and Commission enforcement powers over them began on 2 August 2026. Regulation (EU) 2026/1744 moved the high-risk dates: stand-alone Annex III systems must comply from 2 December 2027, and AI in Annex I regulated products from 2 August 2028.

Find out what your file can and cannot show

Tell us which systems and language claims are in scope and which deadline you are working to. A partner will reply with the questions we need answered to scope a gap assessment.

Request a scoping conversation

Engagements begin under a mutual non-disclosure agreement. Washington, D.C. Serving the United States, the European Union and clients worldwide.

Sources

  1. Regulation (EU) 2024/1689 of the European Parliament and of the Council (Artificial Intelligence Act). Official Journal of the European Union, 12 July 2024. https://eur-lex.europa.eu/eli/reg/2024/1689/oj
  2. Regulation (EU) 2026/1744 (Digital Omnibus on AI). Official Journal, 24 July 2026; in force 27 July 2026. Council of the EU press release, 29 June 2026. https://www.consilium.europa.eu/en/press/press-releases/2026/06/29/artificial-intelligence-council-gives-final-green-light-to-simplify-and-streamline-rules/
  3. California Senate Bill 53 (2025), Transparency in Frontier Artificial Intelligence Act. Business and Professions Code § 22757.10 et seq. https://leginfo.legislature.ca.gov/faces/billNavClient.xhtml?bill_id=202520260SB53
  4. ISO/IEC 42001:2023, Information technology — Artificial intelligence — Management system. https://www.iso.org/standard/81230.html
  5. European Commission, Template for the public summary of training content for general-purpose AI models, 24 July 2025. https://digital-strategy.ec.europa.eu/en/faqs/template-general-purpose-ai-model-providers-summarise-their-training-content
  6. Regulation (EU) 2022/2065 (Digital Services Act), Articles 34 and 35. https://eur-lex.europa.eu/eli/reg/2022/2065/oj
  7. Yong, Z.-X., Menghini, C. and Bach, S. H. (2023). Low-Resource Languages Jailbreak GPT-4. arXiv:2310.02446. https://arxiv.org/abs/2310.02446
  8. Common Crawl Foundation. Statistics of Common Crawl monthly archives: distribution of languages, CC-MAIN-2026-34. https://commoncrawl.github.io/cc-crawl-statistics/plots/languages

Reviewed by the Ariana Nexus Cultural Compliance Bureau. Last updated 19 September 2026.